A new head-to-head trial reveals that while human doctors falter when listing multiple possible diagnoses, artificial intelligence maintains its accuracy.
When a patient rushes into the emergency room, the first guess is rarely the only one that matters. Doctors must quickly build a list of competing possibilities to avoid missing a rare, deadly condition. Yet, a new study shows that this critical skill—differential diagnosis—is exactly where human clinicians struggle, while artificial intelligence holds steady.
This finding challenges the traditional belief that AI is merely a lookup tool for simple tasks. Instead, it suggests that human cognitive fatigue limits our ability to think broadly under pressure, whereas algorithms excel at maintaining an unbiased search space.
The study evaluated 10 emergency medicine physicians against 4 large language models: ChatGPT (GPT-5.2), Gemini 3, Microsoft Copilot (GPT-4), and Claude Opus 4.1. Using 10 standardized clinical vignettes, researchers analyzed 280 diagnostic evaluations. The overall diagnostic accuracy across all tests was 61.79%.
The differential diagnosis gap
The real divide emerged when comparing definitive diagnoses to differential lists. Across all evaluators, definitive accuracy was 70.71%, while differential accuracy dropped to 52.86%.
Human physicians saw their accuracy plummet from 69.0% for definitive diagnoses to a worrying 45.0% for differentials. In stark contrast, the AI models showed no such drop, maintaining 75.0% accuracy for definitive diagnoses and 72.5% for differentials.
This consistency aligns with recent findings on how AI can support clinical decisions. For instance, a study in Computers in Biology and Medicine evaluated large reasoning models as decision support tools in emergency internal medicine, highlighting their utility in complex triage. However, as explored in JAMA Network Open, the actual influence of LLMs on diagnostic reasoning depends heavily on how doctors interact with these suggestions.
Key performance metrics
- AI models achieved 73.75% overall accuracy compared to 57.00% for physicians.
- Physicians experienced a significant decline in accuracy when moving from definitive to differential diagnoses (69.0% down to 45.0%).
- AI models maintained stable performance across both definitive (75.0%) and differential (72.5%) tasks.
Why this matters
This finding reframes the role of AI in the clinic. We often worry about AI “hallucinations” muddying the waters, but human cognitive bandwidth is the actual bottleneck in high-stress emergency settings. If doctors struggle to generate comprehensive differential lists, AI should not just be a passive reference. It must act as an active cognitive safety net to prevent premature diagnostic closure.
A necessary reality check
We must temper these results with the study’s clear limitations. The trial relied on just 10 clinical scenarios and 14 total evaluators, making it a preliminary look rather than a final verdict. Furthermore, these were text-only vignettes, stripped of the chaotic, real-world sensory inputs of an actual emergency department.
Read the full study in the International Journal of Emergency Medicine.
