🧑🏼‍💻 Research - August 21, 2026

AI Predicts Heart Failure From Routine ECGs

🌟 Stay Updated!
Join AI Health Hub to receive the latest insights in health and AI.

An AI model can spot heart failure risk day-by-day using standard electrocardiograms, but its performance drops when crossing oceans.

Doctors can easily diagnose heart failure once the heart has already weakened. The real challenge is catching the decline before the damage becomes permanent. Traditional AI tools only offer a static snapshot, telling clinicians if a patient has a condition right now or predicting risk at one fixed future date.

This new study shifts the strategy from a snapshot to a movie. By analyzing raw 12-lead electrocardiograms, the deep learning model calculates a patient’s risk of developing heart failure with reduced ejection fraction on a day-by-day basis over five years. This continuous tracking could allow clinicians to intervene weeks or months before physical symptoms appear.

Tracking risk over time

Researchers developed and tested the survival model using data from 458,884 patients across three diverse hospital cohorts. The algorithm performed exceptionally well on its home turf. In the development cohort at Zhongshan Hospital, it achieved a C-index of 0.971 (95% CI, 0.965-0.976). When tested at a second hospital in Shanghai, the model maintained a strong C-index of 0.945 (95% CI, 0.938-0.950).

The real test came when the algorithm crossed the ocean to Beth Israel Deaconess Medical Center in Boston. In this US cohort, the model’s accuracy dropped to a C-index of 0.855 (95% CI, 0.850-0.860). While still clinically useful, this performance gap reveals a familiar bottleneck in medical AI.

This drop is where the clinical reality sets in.

Algorithms trained in one healthcare system often struggle when dropped into another. Differences in patient demographics, medical equipment, and clinical workflows can degrade performance. This finding suggests that while a single global model is a noble goal, local calibration remains essential for patient safety.

Opening the clinical black box

Many doctors remain skeptical of deep learning because they cannot see how the algorithm reaches its conclusions. To build trust, the researchers used attention-based frameworks to map the model’s decision-making process. The AI successfully focused on established cardiac markers, including QRS duration, heart rate, and the QT interval.

Here are the key performance metrics from the multinational validation:

  • An internal discrimination score of 0.971 in the development cohort.
  • An external validation score of 0.945 in a neighboring East Asian hospital.
  • A lower but highly competitive score of 0.855 at a major US medical center.
  • Continuous risk tracking mapped across a 5-year horizon.

Despite these strong numbers, the study has limitations. It relies on retrospective data, meaning we do not yet know if using this tool actually improves patient survival. Before clinicians adopt this system, prospective clinical trials must prove that early AI warnings lead to better treatment decisions.

Read the full preprint at medRxiv.

Share on facebook
Facebook
Share on twitter
Twitter
Share on linkedin
LinkedIn
Share on whatsapp
WhatsApp

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.