← Back to AI Health Hub

AI discharge summaries beat human drafts in trial

Automated discharge summaries for long hospital stays reduce dangerous information gaps without increasing patient risk, challenging the assumption that AI cannot handle complex clinical narratives.

Automated discharge summaries for long hospital stays reduce dangerous information gaps without increasing patient risk, challenging the assumption that AI cannot handle complex clinical narratives.

When patients spend weeks in a hospital, their discharge summaries become bloated, messy, and prone to dangerous omissions. Rushed doctors routinely leave out critical follow-up needs. A new study shows that large language models can draft these complex summaries better than the treating physicians, potentially saving hours of administrative work while improving care transitions.

This challenges the prevailing caution that generative AI is too risky for complex, multi-week hospitalizations. For years, skeptics argued that while AI might handle a simple overnight stay, it would fail during a messy 21-day admission. Instead, the data suggests the opposite. Humans are the ones who struggle to synthesize weeks of complex medical events, while algorithms excel at tracking the thread.

This shift builds on earlier evidence, such as a 2025 study in the Journal of Hospital Medicine, which found AI-generated summaries scored higher in continuity of care.

The Trial Results

Researchers evaluated 60 Internal Medicine hospitalizations lasting between 7 and 21 days. Hospitalists and primary care physicians (PCPs) conducted a paired assessment of both human-written and LLM-generated summaries. The performance gap was stark:

  • Clinicians preferred the LLM-generated summaries in 95% of the encounters.
  • The AI drafts had fewer factual inaccuracies and significantly fewer clinically relevant omissions.
  • For 31 incidental radiology findings caught by the AI, 93.5% were factually correct and 87.1% were deemed appropriate to report.
  • There was no significant difference in estimated harm potential or likelihood between the human and AI versions.

The Perception Gap

The most telling finding is not just that the AI won, but who appreciated it. PCPs rated the AI summaries much higher than hospitalists did for understanding and communicating follow-up care. Yet, PCPs also flagged more omissions and assigned a higher likelihood of harm than hospitalists across the board.

This disconnect highlights a systemic friction point in medicine. Hospitalists write summaries to close a chart, while PCPs read them to manage a living patient. AI seems to bridge this gap by default-reporting details that hospitalists dismiss but PCPs desperately need. This aligns with findings from a 2025 study in Communications Medicine on clinical note summarization, which emphasized the need to tailor AI outputs to specific clinical roles.

The Reality Check

This was a retrospective evaluation of just 60 cases. It measures what clinicians prefer on paper, not actual patient outcomes or readmission rates. AI cannot yet run autonomously. A human must still review the draft to catch the remaining 12.9% of inappropriate or incorrect incidental findings.

The practical takeaway is clear. Health systems should not wait for AI to become fully autonomous before deployment. Instead, they should implement LLMs immediately as draft-generation tools to reduce administrative burnout, keeping clinicians strictly in the loop as editors.

Read the full study in npj Digital Medicine.

This article is for informational purposes only and is not a substitute for professional medical advice, diagnosis or treatment.