A new multicenter study reveals that even advanced machine learning models struggle to predict early pregnancy in patients with polycystic ovary syndrome, performing barely better than a coin toss.
Clinicians hoping that machine learning can simplify fertility counseling for patients with polycystic ovary syndrome (PCOS) must temper their expectations. A new study analyzing data from 994 patients across 21 research centers shows that algorithms cannot reliably predict which patients will achieve early pregnancy. The best-performing model achieved an Area Under the Curve (AUC) of just 0.56, a sobering reminder that biological complexity often defeats computational modeling.
This finding directly challenges the prevailing narrative that feeding clinical data into an algorithm yields actionable prognostic tools. While researchers have successfully used machine learning to classify and diagnose the condition itself—as explored in IEEE Access—predicting actual clinical outcomes like pregnancy is a far steeper hill to climb. This performance gap suggests that the active ingredients of fertility cannot be captured by standard clinical metrics alone.
How the models performed
Researchers built three machine learning models—random forest, support vector machine, and decision tree—using a dataset split into a 70% training cohort and a 30% validation cohort. They screened variables using LASSO regression, identifying four key predictors: body mass index (BMI), body weight, endometrial thickness, and sex hormone-binding globulin (SHBG). Across all three models, the predictive efficacy remained low, with AUCs ranging from a weak 0.505 to 0.56.
The random forest model emerged as the top performer, but its metrics highlight the danger of relying on it for clinical decisions:
- An overall accuracy of just 50.8%.
- A sensitivity of 65.1% paired with a low specificity of 47.0%.
- A positive predictive value (PPV) of 24.8% and a negative predictive value (NPV) of 83.3%.
- A low F1 score of 0.360.
This performance is not a validation of AI readiness. It is a warning. Other reviews, such as one in the Eastern Medical College Journal, have noted the potential of AI in early detection, but translating diagnostic accuracy into therapeutic foresight remains an unsolved problem.
The limits of data
These numbers mean the model is highly likely to generate false positives, potentially giving patients false hope about their immediate pregnancy prospects. While the high negative predictive value of 83.3% might help rule out successful early pregnancy in some cases, the overall accuracy makes it too unstable for routine clinical use. This limitation stems from the retrospective nature of the data and the omission of dynamic, real-time lifestyle or environmental factors that influence conception.
For now, clinics should avoid integrating these specific algorithms into patient workflows. Instead of automated prognoses, clinicians must continue to rely on holistic, individualized counseling. The study proves that predicting fertility requires looking beyond basic baseline biomarkers.
Read the full study in the Journal of Ovarian Research.



