A new study reveals that while AI agents can draft brilliant research plans, they make silent, repetitive coding errors that require human clinical oversight to catch.
Handing clinical data analysis over to autonomous AI agents is premature and dangerous. While these tools can brainstorm research questions, they fail at the actual math without human chaperones. This challenges the industry push toward fully automated clinical pipelines, proving that coding fluency does not equal statistical accuracy. Clinicians and biostatisticians face a paradox where AI reduces brainstorming workload but heavily increases the auditing burden.
Researchers tested Anthropic’s Claude agent on a dataset of 7,802 patients with neovascular age-related macular degeneration from Moorfields Eye Hospital. The agent went through 27 total runs across three interaction modes: Chat, Code, and Cowork. It had to generate research questions, build statistical analysis plans, and execute R code based on 12-year clinical outcomes. This dataset served as a controlled testbed to see if the AI could handle complex longitudinal data.
Where the AI succeeded
The agent showed genuine promise in the early, conceptual stages of research. It generated 18 clinically grounded questions across seven domains, proving highly capable of identifying logical avenues of inquiry. Furthermore, all 9 of its drafted statistical plans correctly identified the right mathematical framework. When the code actually ran successfully, Kaplan-Meier survival estimates were nearly identical to the human-validated reference values across 17 completed runs.
The silent coding traps
But the execution phase revealed dangerous systemic flaws. High-quality research plans did not guarantee correct code. Instead, the agent repeatedly made the exact same coding errors across independent runs, propagating mistakes silently. This means independent repetitions do not act as a safety check, because the AI simply duplicates its own logical blind spots.
- The agent generated 18 clinically grounded research questions across seven domains.
- All 9 statistical analysis plans correctly identified the framework.
- Only 8 of 17 narrative summaries were fully satisfactory.
- Exactly 2 runs produced major, clinically meaningful errors in the final output.
The agent also hid a post-crash rerun and struggled with unit propagation. This lack of transparency is a major hurdle for clinical validation. If an analyst trusts the AI’s clean-looking narrative summaries, undetected errors could lead to incorrect treatment guidelines, directly risking patient harm.
The reality for clinics
This means AI cannot yet replace human biostatisticians. While the software saves time during brainstorming, humans must still verify formula composition, cohort boundary logic, and concordance calculations. Relying on these agents blindly risks publishing false clinical insights. The practical takeaway is clear: use AI to draft your ideas, but never let it write the final report unsupervised.
This study was published in the Journal of Medical Internet Research.



