Standard hospital coding misses over half of emergency cases involving drugs, alcohol, and self-harm, but a local AI model can find them.
How can health systems solve a public health crisis if they cannot even count it? For decades, policy decisions and funding for addiction and mental health have relied on standard hospital billing codes. A new study reveals these official records are dangerously incomplete. This leaves hospital administrators blind to the real volume of psychiatric emergencies walking through their doors.
This is not just an administrative headache. It is a systemic failure of data. If half the patients needing help are invisible in the database, resource allocation is essentially guesswork. This study challenges the reliance on traditional administrative databases for health policy. It suggests we are severely underfunding psychiatric emergency services because our tools cannot read the clinical notes.
Researchers analyzed a validation cohort of 2,256 patients aged 16 and older at a UK Type 1 Emergency Department. They compared standard clinical coding against a human-adjudicated reference standard, clinician reviews, and a local large language model (LLM). The reference standard proved that 12.1% of emergency visits actually involved alcohol, drugs, or self-harm. Yet, official clinical coding caught only 6.0% of them, while clinicians flagged 10.0% and the LLM identified 15.6%.
When scaled to 105,096 annual visits, the gap became a chasm. The analysis revealed that 12,890 domain involvements go completely undetected in coded data every year. Traditional coding also failed to capture overlapping issues. It recorded just 1.07 domains per patient, compared to the 1.32 identified by the reference standard.
AI outperforms human review
The LLM did not just beat the administrative codes. It matched or outperformed busy clinicians who reviewed the charts.
- For alcohol cases, the LLM achieved a balanced accuracy of 0.942 compared to 0.930 for clinicians.
- In drug-related cases, the LLM scored 0.959 accuracy, easily beating the clinician score of 0.791.
- For self-harm, the LLM reached 0.982 accuracy, outperforming the clinician baseline of 0.908.
- The model also revealed that 81.6% of self-harm patients required physical medical assessment for injuries or overdoses before any psychiatric review could occur.
That disconnect is the real story.
The clinical reality is complex. Patients rarely present with a single, neat diagnosis. A person might arrive with an overdose that requires immediate physical stabilization, but the underlying driver is self-harm and alcohol abuse. Traditional coding systems force clinicians to pick a primary code, erasing the secondary factors that drive long-term healthcare costs.
The limits of local AI
While these numbers are impressive, we must look at the limitations. This was a single-site study using a locally deployed model. Translating these results to other hospitals with different electronic health record systems remains unproven. Furthermore, the LLM’s higher detection rate of 15.6% suggests it may over-flag some cases, requiring human oversight to prevent false positives.
Relying on manual clinical coding for public health planning is no longer defensible. This study proves that the narrative of the overwhelmed emergency room is even worse than official statistics suggest. Deploying local AI models within existing hospital infrastructure is not a luxury. It is a necessary diagnostic tool for the healthcare system itself.
Read the full study in medRxiv.
