Large language models can draft medical reports, but high error rates mean clinicians spend their saved time correcting the software’s mistakes.
Hospitals hoping to slash radiology backlogs with generative AI are hitting a wall of manual corrections. While algorithms can quickly spit out drafts, they do not yet replace human eyes.
Instead of speeding up workflows, these tools often shift the burden from writing to tedious editing. This reality check comes from a massive systematic review of 101 studies evaluating large language models (LLMs) in medical imaging. The analysis reveals a stark disconnect between technical hype and clinical utility. None of the reviewed studies had a low risk of bias, with 72 rated as high risk and 14 as serious risk.
The safety trade-off
The data shows that AI performance closely mimics human output but carries dangerous blind spots. In a large chest x-ray study, radiologists accepted AI-generated reports 70.5% of the time (6047/8580), compared to a 73.3% acceptance rate (6288/8580) for human-written reports. However, the AI’s false-negative rate was higher, missing critical findings in 18.5% of cases (1584/8580) compared to the human rate of 17.8% (1527/8580).
This margin of error means clinicians cannot trust these tools unsupervised. In another chest x-ray trial, AI reports were equivalent or preferred in 77.7% (233/300) and 56.1% (170/303) of cases across two datasets. Yet, clinically significant errors persisted.
This matches broader industry concerns highlighted in recent literature on multimodal LLMs in medical imaging, which warns that clinical readiness remains a distant goal.
Mixed workflow impacts
The promise of time savings also falls flat under close scrutiny. In a brain MRI study, AI assistance successfully cut initial reading time from 61 to 53 seconds. However, drafting the final impression actually increased overall editing time and edit distance.
Doctors spent more effort rewriting the AI’s awkward phrasing than they would have saved. This pattern shows why we must separate diagnostic accuracy from operational readiness.
As researchers note in a survey on large language models in medical image analysis, generating readable text is not the same as generating safe, clinically useful text.
The path forward
Here is what the data actually proves:
- AI report acceptance reached 70.5% but came with a 18.5% false-negative rate in chest x-rays.
- Brain MRI reading times dropped by 8 seconds, but editing workloads increased.
- Out of 101 studies, zero were judged to have a low risk of bias.
The practical takeaway is clear. Healthcare systems should not deploy these tools autonomously. Until developers standardize how they report omissions, commissions, and failed generations, AI drafting should remain strictly assistive and locally validated.
Read the full analysis in the Journal of Medical Internet Research.



