A new validation study reveals that while AI can safely rule out normal chest scans to ease radiologist workloads, severe software calibration issues could trigger unnecessary follow-up tests.
An AI tool that flags abnormal lung scans with high accuracy sounds like an instant win for overworked clinics. But when deployed in Nepal, a prominent deep learning model successfully sorted the healthy from the sick while suffering from a hidden technical glitch: severe score compression. This means the AI got the answers right but with dangerously narrow confidence formatting, illustrating why off-the-shelf medical software cannot be trusted without local tuning.
The operational reality is a mixed bag. For the 824 health assessment applicants screened at Patan Hospital, the DenseNet-121 algorithm achieved a high negative predictive value of 99.67%. In practice, this means clinicians could use the tool to safely bypass manual reviews for hundreds of normal scans. This would dramatically cut down waiting times for migration and student visas.
However, the positive predictive value was a dismal 17.89%. This low rate means that for every five patients the AI flagged as abnormal, four were actually healthy. In a busy migration screening clinic, this bottleneck would overwhelm secondary testing pipelines and cause needless patient anxiety.
Performance by the numbers
- The algorithm flagged abnormal scans with 95.12% sensitivity and 77.2% specificity.
- Out of 824 completed cases, only 4.97% (41 scans) were flagged as truly abnormal by radiologists.
- The model’s expected calibration error was 0.564, showing extreme score compression between 0.52 and 0.72.
- The low Cohen’s kappa of 0.237 highlights poor overall agreement with human readers on exact classifications.
The calibration trap
The real warning here lies in the calibration failure. The algorithm compressed its risk scores into a narrow band, meaning it could not reliably tell a slightly abnormal scan from a highly critical one. This is a classic symptom of geographic dataset shift, where an AI trained on Western data struggles with the demographics of a South Asian hospital.
Furthermore, the study only captured 41 true positive cases over its 16-day run. This small sample size leaves a wide 14.8 percentage point uncertainty margin around the sensitivity rate. Crucially, a normal AI triage result does not rule out active pulmonary tuberculosis, which still requires microbiological testing.
Hospitals in low- and middle-income countries should view this as a yellow light. The AI works well as a basic filter, but deploying it without local calibration will lead to massive clinical inefficiency.
Read the full study in BMJ Open.



