← Back to AI Health Hub

Cheap AI models beat specialized clinical software

Hospitals trying to automate lung cancer screening eligibility no longer have to choose between speed, cost, and clinical accuracy.

Hospitals trying to automate lung cancer screening eligibility no longer have to choose between speed, cost, and clinical accuracy.

For years, the rule of thumb in clinical NLP was simple. If you wanted fast, cheap, and reliable data extraction, you built a lightweight, specialized model. Generative AI was considered too slow, too expensive, and too prone to making up math.

A new benchmark study flips this logic on its head. General-purpose large language models are now beating specialized clinical software on both cost and reasoning. This shift challenges how health systems should build their clinical decision support pipelines.

Testing the models

Researchers tested this shift using a synthetic benchmark of 3,000 outpatient notes split into three distinct scenarios. They compared a dedicated structured-judgment model, TypeSafe Jev 1.13, against four major language models: Claude Haiku 4.5, Claude Sonnet 5, GPT-6 Luna, and GPT-6 Sol. The goal was to extract smoking status, pack-years, and quit dates to determine lung cancer screening eligibility.

The results show a clear trade-off between raw speed and mathematical accuracy.

  • TypeSafe Jev 1.13 achieved eligibility accuracy of 99.1% on template notes, 99.9% on messy notes, but fell to 94.4% on notes requiring complex math.
  • GPT-6 Luna reached 99.7%, 98.8%, and 98.5% accuracy across the same three categories.
  • Claude Sonnet 5 was the most accurate, hitting 100.0%, 99.6%, and 99.8% accuracy.
  • In the complex math category, Jev generated 27 false positives and 14 false negatives per 1,000 notes, the worst performance in the cohort.
  • Jev was the fastest model, processing notes in 0.45 to 1.21 seconds, compared to 1.71 to 3.50 seconds for the other models.
  • GPT-6 Luna was the cheapest option, costing just $0.12 to $0.21 per 1,000 notes, while Jev cost $0.61 to $0.64 and Sonnet cost up to $6.61.

Rethinking clinical pipelines

This data complicates the standard playbook for clinical IT departments. Jev is fast enough for real-time, synchronous clinical decision support. Doctors do not have to wait for batched processing during a patient visit.

Yet Jev’s math failures are a liability. Missing a patient’s screening eligibility because of a poor pack-year calculation is a direct clinical risk. Meanwhile, GPT-6 Luna destroys the argument that general models are too expensive for bulk clinical tasks. At less than a third of Jev’s cost, Luna offers superior accuracy on complex notes, though its higher latency remains a bottleneck for real-time use.

This means developers must choose between immediate feedback and clinical safety. If a system requires instant alerts while a doctor is typing, the specialized model wins on speed. But if accuracy in calculating pack-years is the priority to avoid missed diagnoses, the general models are the only viable choice.

The study is limited by its use of synthetic notes. Real-world clinical records contain even messier formatting and unstructured jargon that could degrade performance across all models.

For health systems, the practical takeaway is clear. Stop building bespoke extraction tools when cheap, general-purpose models can do the math better for a fraction of the price.

Read the full study in medRxiv.

This article is for informational purposes only and is not a substitute for professional medical advice, diagnosis or treatment.