Medical students using Chinese reasoning model DeepSeek R1 scored significantly higher on urology exams than those using ChatGPT o3-mini or traditional web search.
Medical schools are scrambling to integrate artificial intelligence into their curricula without clear evidence of which tools actually improve learning. A new randomized controlled trial published in the Journal of Medical Internet Research shows that model choice directly dictates educational success. The findings challenge the assumption that market-leading American models are automatically superior for specialized medical training.
This dynamic highlights a shift in digital medical pedagogy. While early research focused on how large language models in medical education could assist faculty, this trial demonstrates that specific reasoning architectures alter student performance directly during independent study.
Testing accuracy and student scores
Researchers evaluated the accuracy of two reasoning models, ChatGPT o3-mini and DeepSeek R1, using authoritative urology multiple-choice questions. They then ran a randomized controlled trial comparing undergraduates using DeepSeek R1, ChatGPT o3-mini, or traditional internet-based self-study methods. Post-test results and student feedback were collected to measure practical learning gains.
- DeepSeek R1 demonstrated higher accuracy than ChatGPT o3-mini on benchmark urology questions.
- The DeepSeek R1 group outperformed both traditional internet searchers and ChatGPT o3-mini users in total post-test scores across multiple question types.
- ChatGPT o3-mini produced higher numerical scores than traditional study methods, but the difference failed to reach statistical significance.
- Survey data revealed that most participating medical undergraduates held a positive attitude toward AI-assisted learning.
The gap between the two models matters for institutional deployment. DeepSeek R1 successfully converted self-study into measurable exam performance, whereas ChatGPT o3-mini performed no better than standard web browsing in a statistically meaningful way.
That variance changes how medical educators should evaluate digital study aids.
When selecting AI tools, schools cannot simply rely on brand recognition or general benchmark performance. Specialized medical recall and reasoning vary widely between models, meaning an unvalidated tool could waste student time without improving clinical knowledge retention. Scholars analyzing large language models in worldwide medical exams have noted similar performance gaps across varied platforms.
Study limitations and real-world impact
The trial carries clear operational limitations. The study evaluated short-term test scores in a single surgical subspecialty, urology, rather than long-term clinical retention or hands-on patient care. Multiple-choice accuracy does not guarantee that a student can manage real-world clinical uncertainty or apply judgment at the bedside.
Medical institutions should avoid blanket rollouts of consumer AI tools. Instead, faculty must run subject-specific validations before recommending particular models to students.
Read the full peer-reviewed trial in the Journal of Medical Internet Research.



