A simple change to multiple-choice questions reveals that medical AI models may rely on pattern recognition rather than actual clinical reasoning.
What happens when you ask a top-tier medical AI a question, but swap the correct option for “None of the other answers”? It turns out the systems we trust to pass medical exams might just be guessing based on visual patterns. This simple swap exposes a massive gap between memorizing a textbook and understanding a patient.
For years, AI developers have boasted about models passing board exams with flying colors. This study suggests those high scores are an illusion. By relying on multiple-choice formats, we are measuring test-taking tricks rather than robust clinical competence. This disconnect is the real story.
Testing the models
Researchers took 210 board renewal questions from the Japanese Society of Nephrology spanning 2014 to 2023. Two independent nephrologists reviewed the items and modified them, leaving 145 validated questions where “None of the other answers” (NOTA) became the only correct choice. They tested four major models: GPT-5, GPT-4o, Gemini 2.5 Pro, and Gemini 2.0 Flash. This methodology, also discussed in a related nephrology robustness report, highlights how easily AI reasoning chains break under minor pressure.
A steep drop-off
The results were uniform and severe across every single model tested. When forced to evaluate the NOTA option, accuracy plummeted.
- GPT-4o crashed from 66.21% to 19.31%, a massive 46.90 percentage point drop.
- GPT-5 fell from 87.59% to 73.10%, showing the most resilience with a 14.48 percentage point decline.
- Gemini 2.5 Pro dropped from 86.90% to 55.86%, a loss of 31.03 percentage points.
- Gemini 2.0 Flash slid from 58.62% to 31.03%, losing 27.59 percentage points.
Every drop was statistically significant with a p-value of less than .001. While GPT-5 proved to be the most stable, even the best model faltered when the safety net of familiar multiple-choice structures was removed.
Rethinking medical AI benchmarks
This collapse in performance is a warning sign for clinical deployment. Real-world medicine is entirely open-ended. If an AI relies on the presence of specific distractor answers to find the right path, it will fail when facing a complex kidney patient. We are currently certifying software that cannot handle basic negative constraints.
We must acknowledge the limitations of this data. The study relied on a relatively small set of 145 translated Japanese board questions, and preprints like this require further peer review. However, the core lesson is clear. We must stop using standard multiple-choice tests as the gold standard for AI safety. Until models can handle negative constraints, they remain advanced search engines, not digital doctors.
Read the full preprint on medRxiv.



