🧑🏼‍💻 Research - September 3, 2026

AI did not improve doctors’ diagnostic accuracy

🌟 Stay Updated!
Join AI Health Hub to receive the latest insights in health and AI.

A new clinical trial reveals that certified AI diagnostic software saves doctors time but fails to make their diagnoses any more accurate.

Why do we assume faster medical decisions are better decisions? If a digital assistant cuts a doctor’s research time in half but leaves them just as likely to misdiagnose a rare disease, we have not solved the clinical bottleneck. We have only accelerated the rate of confident mistakes.

This trial challenges the assumption that certified clinical large language models are ready to act as safety nets. The findings show that while AI feels like a helpful partner, it does not actually improve clinical performance. This suggests that the current push toward AI integration in complex specialties may be built on a psychological illusion of safety.

The clinical test

The ALLIANCE trial evaluated 82 physicians across seven hospitals in two countries. Researchers split the doctors into two groups to diagnose complex rheumatology cases. One group used conventional search tools, while the other had access to Prof. Valmed, a certified diagnostic support system.

The results challenge the optimistic outlook of earlier papers, such as a 2025 study on hybrid care and task shifting in rheumatology. Both groups of doctors improved after consulting their assigned tools, but the AI group showed no unique advantage.

Speed without accuracy

The data paints a stark picture of efficiency decoupled from accuracy:

  • Top-1 diagnostic accuracy rose from 22.2% to 33.3% with AI, compared to a nearly identical rise from 23.3% to 35.0% using standard tools (adjusted OR 0.99, p=0.979).
  • AI-assisted doctors worked much faster, taking just 94 seconds per case compared to 206 seconds for the control group (adjusted mean difference of -112 seconds, p<0.001).
  • There were no significant differences in top-3 accuracy, diagnostic reasoning, or doctor confidence between the two groups.

This speed-accuracy trade-off is particularly concerning given how well these models perform in isolation. Benchmarks show they can score highly on standardized tests, as documented in a 2025 evaluation of rheumatology multiple-choice questions. But when integrated into real clinical workflows, that raw capability fails to translate into better human decisions.

The over-reliance trap

The real danger lies in how doctors perceive these systems. Physicians using the AI rated the timeliness and quality of the support much higher, despite getting no objective diagnostic boost. The researchers also flagged persistent overconfidence and a worrying level of AI over-reliance among participants.

This creates a dangerous feedback loop. Doctors get answers faster, feel more supported, and trust the output, yet they are no more accurate than colleagues using basic search engines. If we deploy these tools widely, we risk codifying a system of rapid, highly confident errors.

We must stop evaluating clinical AI solely on its standalone accuracy or its ability to save time.

Read the full study in medRxiv.

Share on facebook
Facebook
Share on twitter
Twitter
Share on linkedin
LinkedIn
Share on whatsapp
WhatsApp

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.