A new study reveals that training AI to ace medical benchmarks actually makes its clinical reasoning worse, exposing a dangerous gap in how we build healthcare algorithms.
Why does an AI that scores near-perfect on paper feel completely wrong to a practicing doctor?
We often assume that optimizing an algorithm’s math automatically makes it a safer clinical partner. A peer-reviewed study challenges this assumption, revealing what researchers call an “alignment paradox” in fertility care. When algorithms are tuned to chase perfect metric scores, they lose the nuanced reasoning that keeps patients safe.
This disconnect should force a complete rethink of how we validate medical AI. If high benchmark scores do not translate to better clinical decisions, then our current testing standards are functionally useless.
The Illusion of Progress
Researchers built four alignment strategies on a shared open-source biomedical backbone. They tested them using 8,201 deidentified electronic health records from West China Second University Hospital, collected between January 2020 and December 2022, with a patient mean age of 31.79 years.
On paper, the advanced Group Relative Policy Optimization (GRPO) model looked like the clear winner. It achieved the highest average automatic performance across five structured decision fields.
The top-performing GRPO model achieved the following results on key metrics:
- Infertility type accuracy reached 92.57%.
- ART strategy accuracy hit 76.49%.
- COS regimen accuracy landed at 62.36%.
- Gonadotropin starting dose mean absolute error was 44.94.
But these numbers hide a deeper failure. When two reproductive medicine specialists ran a blinded review of 100 paired cases, the simpler, conservative Supervised Fine-Tuning (SFT) baseline actually outperformed the advanced GRPO model in clinical utility.
Why Math Fails Medicine
In a three-way best-response comparison, experts chose the SFT baseline as the best response in 51.2% of cases. GRPO was selected in only 26.2% of cases, while the original physician-charted plan was chosen 22.6% of the time. This does not mean the AI is better than human doctors, but it shows that SFT aligned far better with expert clinical reasoning than the mathematically superior GRPO.
The reason for this split is simple. Advanced optimization models reward token-level correctness but ignore the holistic reasoning clinicians rely on. For instance, GRPO improved scores in standard IVF and preimplantation genetic testing (PGT) cases, but its performance fell in intracytoplasmic sperm injection (ICSI) cases because it failed to parse unstructured male-factor clinical data.
Worse, both models kept making things up. The hallucination rate was 15% for GRPO and 18.5% for SFT. That means nearly one in five recommendations contained clinically unsupported content.
The Path Forward
This study exposes a hard truth: we cannot benchmark our way to safe medical AI. If a model can ace a test while hallucinating dangerous drug doses, the test itself is broken.
We must be honest about the limitations of this data. The clinical review relied on only two specialists who had marginal interrater agreement, and the study was a retrospective, single-center trial.
Until we validate these tools across multiple centers with diverse patient populations, neither model is ready for clinical deployment. Developers must stop chasing abstract benchmarks and start measuring what actually happens when a doctor sits down with a patient.
Read the full study in the Journal of Medical Internet Research.



