← Back to AI Health Hub

AI vessel segmentation made radiologist diagnoses worse

A highly accurate deep learning tool designed to help radiologists spot vertebral artery tears actually degraded their diagnostic performance.

A highly accurate deep learning tool designed to help radiologists spot vertebral artery tears actually degraded their diagnostic performance.

Why do we assume that cleaner medical images lead to better clinical decisions? A new study on vertebral artery dissection (VAD) exposes a dangerous disconnect between AI engineering and human psychology. The algorithm did its job near-perfectly, yet the doctors reading the scans made more mistakes.

This challenges the industry’s obsession with pixel-level accuracy. We are building beautiful maps that lead drivers off the cliff. It is time to rethink how we validate diagnostic software.

The diagnostic paradox

Between September 2024 and July 2026, researchers developed a deep learning model using nnU-Net version 2 to segment vascular structures on CT angiograms. They trained the model using five-fold cross-validation on 101 CTAs, combining 84 manually segmented internal scans with 17 from the RSNA Intracranial Aneurysm Challenge. Testing took place on an external cohort of 40 CTAs, which included 22 positive and 18 negative cases of vertebral artery dissection.

On a technical level, the AI was a triumph. It achieved a Dice similarity coefficient of 0.96 (SD 0.01) out of 1.0, proving it could outline blood vessels with extreme precision.

The real test came when 10 human readers evaluated the scans with and without the AI’s visual assistance. Instead of sharpening their eyes, the digital overlays clouded their judgment. The paired crossover study revealed stark outcomes:

  • Diagnostic accuracy dropped significantly from 0.84 without the AI to 0.72 with it (P < 0.001).
  • Interpretation speed did not improve, shifting insignificantly from 166.7 seconds to 137.3 seconds (P = .54).
  • Subjective confidence remained flat, moving from 3.96 to 3.98 (P = .77).

Yet, the most alarming finding is subjective. Despite performing worse with the tool, 7 out of 10 readers reported they would still routinely use it in acute or trauma settings. This reveals a dangerous cognitive bias. Radiologists may favor the aesthetic comfort of highlighted vessels over their own clinical accuracy.

The danger of false trust

Vertebral artery dissection is a sneaky, high-stakes diagnosis. Missing a tear in these neck arteries can lead to a stroke, making diagnostic accuracy in trauma bays a matter of life or death. The clinical stakes are simply too high for tools that degrade human performance.

This study shows that highlighting a blood vessel does not help a radiologist see a tear inside it. In fact, the bright digital overlays might distract from subtle structural abnormalities, creating a false sense of security. If clinicians trust the tool enough to adopt it despite worse outcomes, we risk scaling systematic diagnostic errors across emergency departments.

The study is limited by its retrospective design and a small external testing set of just 40 scans. However, it serves as a warning. Perfect segmentation does not equal clinical utility. Until developers test how AI overlays affect human decision-making in real-time, these tools should remain out of the reading room.

Read the full preprint in medRxiv.

This article is for informational purposes only and is not a substitute for professional medical advice, diagnosis or treatment.