← Back to AI Health Hub

Structural, Functional and Cognitive Validation of a Deep-Learning MCI-to-AD Conversion Model in OASIS-3

An AI trained to predict Alzheimer's progression successfully spotted high-risk patients in a new dataset, but its absolute risk scores proved dangerously misleading.

AI Alzheimer’s model fails calibration in new test

An AI trained to predict Alzheimer’s progression successfully spotted high-risk patients in a new dataset, but its absolute risk scores proved dangerously misleading.

Can we trust a clinical AI tool when it moves to a different hospital? A new validation study of TAF-Net, a deep-learning model designed to predict the transition from mild cognitive impairment to Alzheimer’s, reveals a troubling disconnect. The algorithm successfully ranks who is at higher risk, but its absolute probability scores are dangerously off.

This is not just a technical glitch. If a clinician relies on this tool today, they might tell a patient they have a near-zero chance of developing Alzheimer’s when their actual risk is nearly one-in-six. It proves that high discrimination scores on paper do not guarantee clinical safety in the real world.

Researchers tested the model, which was originally trained on the ADNI dataset, on **101** patients from the OASIS-3 database without retraining it. This cohort included **25** patients who converted to Alzheimer’s and **76** who remained stable. To verify the biology behind the AI’s math, the team analyzed structural brain scans from **41** participants and functional connectivity data from **40** participants.

Key validation metrics

  • The model achieved an AUC of **0.72** in the new patient cohort.
  • Patients sorted into the lowest risk tiers still converted at a **15%** rate.
  • AI risk scores correlated with reduced functional connectivity in salience and default-mode hubs.

The calibration trap

The model proved it has a grasp on the biological reality of the disease. High risk scores tracked faster tissue loss in the medial-temporal lobe and matched cognitive decline on standard tests. However, the math broke down when calculating actual probabilities.

The model assigned a near-zero probability of conversion to the lowest two-thirds of the cohort. In reality, **15%** of those patients converted to Alzheimer’s. This means the model’s rank-ordering ability transferred to the new hospital dataset, but its calibration failed entirely.

A mirror for atrophy

There is another catch for clinicians hoping for new insights. The study found that the AI’s predictive power was statistically accounted for by native structural atrophy. The algorithm is not finding hidden patterns. It is simply re-expressing the physical brain shrinkage we can already see on a standard MRI.

For AI to survive in clinical practice, developers must look beyond simple discrimination metrics. An AUC of **0.72** looks respectable in a journal, but poor calibration makes it a liability at the bedside. Before these tools can be trusted, they must be recalibrated for the specific populations they are meant to treat.

This study was originally published in medRxiv.

This article is for informational purposes only and is not a substitute for professional medical advice, diagnosis or treatment.