An AI trained on autopsy text can spot rare brain diseases with surprising accuracy, but it fails completely when pathologies mix.
Can a computer read a pathologist’s handwritten notes and know exactly how a patient died? For decades, the final word on neurodegenerative disease has belonged to post-mortem tissue analysis. Yet thousands of these highly detailed autopsy reports sit idle in archives, written in free-form narrative text that traditional software cannot parse.
This study challenges the idea that we need complex molecular testing to categorize these diseases. By feeding 25 years of autopsy notes to large language models, researchers proved that the naked-eye observations of a pathologist contain highly structured diagnostic signals.
This is not just about automating old paperwork. It forces us to rethink what makes a disease distinct. The AI excelled at identifying conditions with clear, localized physical footprints, like progressive supranuclear palsy. But it stumbled badly on mixed cases.
This gap reveals a hard truth for digital pathology. Algorithms are excellent at recognizing textbook, isolated anatomical damage. However, when multiple diseases overlap in the same brain, macroscopic text descriptions lose their utility. We still cannot escape the need for expensive microscopic or molecular verification.
Mining the brain bank
Researchers analyzed 5,613 autopsy cases from the Mayo Clinic Brain Bank collected between 1998 and 2023. The cohort spanned seven major diagnostic categories, including Alzheimer’s, Lewy body disease, and frontotemporal lobar degeneration. A fine-tuned language model read the narrative gross descriptions to extract semi-quantitative scores for 39 anatomical features, achieving an extraction accuracy of 0.95 during manual testing.
How the models performed
The researchers built two classifiers to predict the final neuropathological diagnosis: a CatBoost model using the extracted scores, and a text-based LLM. Both models factored in age at death, sex, and brain weight.
- The text-based LLM achieved 0.75 accuracy and a Cohen’s kappa of 0.68.
- The CatBoost classifier reached 0.73 accuracy and a macro-average AUC of 0.92.
- Both models struggled with mixed pathology, yielding an AD-LBD sensitivity of just 0.21 for CatBoost and 0.03 for the LLM.
- In contrast, PSP sensitivity reached 0.93, and MSA sensitivity reached 0.92.
The limits of anatomy
The stark contrast in sensitivity shows that macroscopic brain structure has strict diagnostic limits. Subthalamic nucleus atrophy and putaminal abnormalities leave obvious physical scars that the AI easily flags. But when Alzheimer’s and Lewy body disease co-exist, they do not create a unique physical signature that a pathologist can see with the naked eye.
For researchers, this means LLMs can rapidly clean and structure massive historical archives of brain bank data. But for clinical diagnostics, it proves that macroscopic AI screening has a ceiling. We cannot rely on structural changes alone to untangle the complex, overlapping pathologies that define dementia in an aging population.
Read the full study in medRxiv.



