Yesil Science YesilScience
Let's talk
Menu

Hospital AI Tool Hallucinates Twelve Percent of Time

A new study reveals that while doctors find EHR-integrated generative AI easy to use, the software regularly invents clinical details and ignores the vast majority of patient records.

hospital ai tool hallucinates twelve percent of time

A new study reveals that while doctors find EHR-integrated generative AI easy to use, the software regularly invents clinical details and ignores the vast majority of patient records.

How much of a patient’s history can an AI ignore before it becomes dangerous? A new evaluation of Epic IP Insights across a seven-hospital health system reveals a stark disconnect between user convenience and clinical accuracy. The tool generated 706 summaries across 445 encounters, but it only analyzed a mean of 5.8% of the available patient notes.

This selective reading compromises the very purpose of automated summarization. If an AI overlooks more than 94% of the clinical record, it is not summarizing the patient’s history. It is cherry-picking it.

The Hallucination Rate

To measure accuracy, researchers built an agentic hallucination detector that broke summaries down into atomic content units (ACUs). The detector agreed with physician adjudication on 96.95% of ACUs in a validation set of 5 summaries (295 ACUs), after being developed on 30 summaries (2,012 ACUs). When applied to the remaining 671 summaries containing 40,452 ACUs, the detector flagged a hallucination rate of 11.79%.

This means nearly one in eight clinical facts presented by the AI was fabricated. Furthermore, the tool lacks consistency. Among 385 regenerated pairs of summaries, only 32.7% were textually identical and only 37.9% were semantically identical.

The Trust Gap

This inconsistency explains why clinicians remain deeply skeptical of the technology. While 97.8% of surveyed clinicians found the tool easy to use, only 45.9% felt comfortable relying on the output with little verification. Only 45.3% expected the tool to yield efficiency gains, and just 46.1% reported using it frequently.

These numbers challenge the industry assumption that ease of use translates to clinical utility. Doctors are willing to click a button, but they are not willing to risk patient safety on unverified AI outputs. The technology currently acts as a cognitive speed bump rather than a timesaver, requiring clinicians to double-check every claim against the source documents.

Why This Matters

This finding forces a rethink of how hospitals adopt generative AI. We cannot treat clinical summarizers like search engines that can afford occasional errors. When an AI hallucinates nearly 12% of the time, the burden of proof shifts entirely to the human operator, defeating the goal of reducing clinician burnout.

Future implementations must move away from simple pilot programs toward rigorous, multidimensional assurance frameworks. Until these tools can reliably ingest the entirety of a patient’s record without inventing symptoms, they remain high-risk assistants that require constant supervision.

Key Evaluation Findings

  • The AI tool ignored more than 94% of available patient notes, utilizing a mean of just 5.8%.
  • The automated detector identified a hallucination rate of 11.79% across 40,452 analyzed clinical units.
  • Fewer than half of clinicians (45.3%) expected the tool to deliver actual efficiency gains.
  • Regenerated summaries matched semantically only 37.9% of the time, showing high inconsistency.

This analysis is based on a study published in medRxiv.