A new benchmark reveals that medical artificial intelligence aces multiple-choice tests but stumbles when faced with real patient cases.
Can an AI that scores an A-grade on a medical licensing exam actually treat a patient? Tech companies want us to believe the answer is yes. But a new study exposes a massive gap between test-taking and actual clinical reasoning.
This gap challenges the current trend of validating AI using standardized multiple-choice tests. It suggests we are measuring the wrong things. We are building elite test-takers, not digital clinicians.
The test-taking illusion
Researchers built the Bones and Joints (B&J) Benchmark using 1245 real-world patient cases from orthopedics and sports medicine. They tested 14 vision-language models (VLMs) and 6 large language models (LLMs) across seven core clinical tasks. These tasks mirrored the actual clinical pathway, from interpreting images to planning treatment.
The results expose a stark divide. On structured, multiple-choice questions, the best models scored above 90% accuracy. But when forced to handle open-ended tasks that required combining text and images, their accuracy plummeted to barely 60%.
This drop-off is not just a minor bug. It proves that high scores on medical licensing exams are an illusion of competence. Models are memorizing patterns, not reasoning. This aligns with broader industry concerns about trust and interpretability in medical imaging, as documented in a recent IEEE Access survey on multimodal models.
Hallucinations in the clinic
The failure is most acute where text meets imagery. The VLMs struggled to interpret medical images, often falling back on text-driven hallucinations. Instead of looking at the X-ray, the AI “guessed” what should be there based on the patient’s written chart.
Surprisingly, specialized medical AI models performed no better than general-purpose models. Buying a specialized clinical model does not buy you safety. To fix this, researchers are exploring new training methods like key concept learning to help models actually reason through visual data rather than just matching text patterns.
- Models scored over 90% on multiple-choice questions but fell to 60% on open-ended clinical tasks.
- Medical-specific models showed no consistent advantage over general-purpose models.
- AI frequently suffered from text-driven hallucinations, ignoring visual evidence in X-rays.
Why this matters
If an AI cannot reliably connect an X-ray of a fractured bone to a treatment plan, it cannot be trusted with independent clinical decisions. Using these models for automated triage or unsupervised diagnostics in orthopedics is currently dangerous.
For now, the safe play is to keep these models in strictly supportive, text-only roles like drafting clinical summaries. True clinical autonomy requires a fundamental breakthrough in how machines integrate sight and language.
This analysis is based on research published in npj Digital Medicine.
