AI spots rare diseases years before doctors do
An algorithm scanning millions of medical records proved that rare genetic conditions leave clear digital breadcrumbs years before a formal diagnosis.
How long should a patient suffer before a doctor connects the dots? For people with rare genetic diseases, the diagnostic odyssey often drags on for a decade. A new preprint suggests this delay is not caused by a lack of data, but by a failure to read it.
This finding challenges the assumption that we need expensive genetic sequencing upfront to flag rare diseases. Instead, the raw material is already sitting unused in standard billing codes and clinical notes. By mining existing electronic health records, health systems can bypass the clinical blindspots that keep rare diseases hidden.
The diagnostic lag
Researchers trained machine learning models on a massive longitudinal cohort of roughly 3 million patient records from the Mayo Clinic. They targeted eight rare conditions, extracting billing codes and phenotypes from clinical notes. The algorithm’s accuracy varied by condition, identifying between 10% and 89% of patients early, while maintaining a strict 99% specificity to avoid overwhelming doctors with false alarms.
The lead times are where the clinical value becomes undeniable. The model flagged patients years before their first official diagnostic code, including:
- 2,755 days early for hereditary angioedema
- 2,320 days early for hereditary hemorrhagic telangiectasia
- 1,408 days early for Fabry disease
- 1,067 days early for neurofibromatosis type 1
This builds on previous efforts, such as a 2024 study in the Orphanet Journal of Rare Diseases, which used supervised machine learning and semantic similarity to detect rare ciliopathy patients. While that work proved the concept in specific niches, this new Mayo Clinic data shows the sheer scale of what is possible across diverse conditions.
The reality check
There is a catch. A 10% detection rate for some conditions means many patients will still slip through the cracks. Moreover, electronic health record data is notoriously messy and incomplete outside of elite institutions like the Mayo Clinic.
If healthcare systems want to deploy these models, they must accept that algorithms are only as good as the clinical notes they feed on. The real bottleneck is no longer technology, but our fragmented record-keeping.
Read the full study in medRxiv.



