🧑🏼‍💻 Research - August 26, 2026

Transformer-Based Survival Model for Cardiovascular Risk Prediction from Longitudinal Health Checkup Data

🌟 Stay Updated!
Join AI Health Hub to receive the latest insights in health and AI.

AI Transformer Predicts Heart Disease From Checkups

A new deep learning model proves that standard health checkups hold hidden, complex patterns that traditional risk calculators completely miss.

For decades, clinics have relied on static scoring systems to predict heart attacks and strokes. These tools treat risk like a simple checklist. But human health is a messy web of habits and biology that changes over time.

This research challenges the status quo of cardiovascular forecasting. By applying a Transformer model to routine health checkups, researchers proved that deep learning can capture complex, non-linear relationships that traditional math ignores. It suggests we no longer need expensive, specialized testing to get highly accurate, personalized risk assessments.

The era of static risk calculators is ending.

The study analyzed a massive development cohort of 100,056 participants without baseline cardiovascular disease from 2010 to 2024. Over a 10-year follow-up, 4,113 cardiovascular events (a 4.1% rate) occurred, defined as self-reported heart disease or stroke. The researchers trained their Transformer model on basic laboratory, anthropometric, and self-reported lifestyle data.

To prove the model works in the real world, they tested it on an external cohort of 79,756 Kanazawa City participants, who had a much higher event rate of 26.6% (21,179 events). The AI went head-to-head with Cox regression, XGBoost, multilayer perceptrons, and classic clinical tools like the Framingham and Hisayama Risk Scores.

The Performance Gap

  • The Transformer achieved an internal 10-year ROC-AUC of 0.821 (95% CI, 0.816–0.826) and a concordance index of 0.781 (CI, 0.775–0.787).
  • External validation remained strong, yielding an ROC-AUC of 0.762 and a C-index of 0.744.
  • The model’s internal precision-recall AUC reached 0.427 (CI, 0.419–0.435), while the external cohort scored 0.500.

Peeking Inside the Black Box

Clinicians often distrust deep learning because of its lack of transparency. To solve this, the researchers used SHapley Additive exPlanations (SHAP) and a Feature-level Attention Network (FAN) to visualize the top 12 SHAP-ranked features. They found that age, electrocardiogram abnormalities, sex, and blood pressure medication were the heaviest anchors of risk.

More importantly, the attention network showed exactly how behavior modifies biology. For instance, daily exercise and weight gain directly shifted how the model calculated age-related risk. Within this network, age acted as a central hub, linking lifestyle choices to physical biomarkers.

However, we must acknowledge a major limitation. The study relied on self-reported physician diagnoses for cardiovascular events, which introduces recall bias. Additionally, the external cohort had a vastly higher disease rate of 26.6% compared to the development cohort’s 4.1%, suggesting the model still needs calibration across different clinical environments.

Why This Matters

This finding matters because it elevates the clinical value of cheap, self-reported lifestyle data. Usually, doctors dismiss patient questionnaires as soft science. This model proves that when processed through an attention-based network, simple habits like daily exercise contain hard predictive power that rivals laboratory blood tests.

Read the full preprint study in medRxiv.

Share on facebook
Facebook
Share on twitter
Twitter
Share on linkedin
LinkedIn
Share on whatsapp
WhatsApp

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.