🧑🏼‍💻 Research - August 10, 2026

A Multimodal large language model-based triage tool for osteoporotic vertebral compression fractures using posture and movement videos

🌟 Stay Updated!
Join AI Health Hub to receive the latest insights in health and AI.

AI spots spinal fractures from smartphone videos

A new AI model screens for painful spinal fractures using simple home videos instead of immediate, expensive hospital scans.

How do you diagnose a broken spine when the patient just thinks they have a bad back?

For years, diagnosing osteoporotic vertebral compression fractures (OVCF) has required patients to get X-rays or MRIs. This creates a massive bottleneck. Many patients simply suffer at home, unaware their spine is collapsing. This new study challenges the assumption that we need immediate imaging to catch these fractures. By turning a smartphone camera into a clinical triage tool, researchers are proving that how a person moves can be just as telling as an expensive scan.

From video to diagnosis

The system uses a two-stage framework. First, a multimodal large language model analyzes posture images and movement videos using pose estimation. It scores the patient on alignment, symmetry, coordination, and pain response. Then, a machine-learning classifier combines these scores with basic clinical data to make the final call.

The researchers trained the system on a multicenter cohort of 204 participants, split evenly between 102 OVCF patients and 102 healthy controls. To prove the tool works in the real world, they tested it on a geographically independent external validation group of 56 patients. The Gradient Boosting classifier performed remarkably well under pressure.

The performance breakdown

The model proved it could handle patients from clinics it had never encountered before. The key metrics show its potential as a gatekeeper for hospital imaging:

  • The model achieved an area under the receiver operating characteristic curve (AUC) of 0.838.
  • It caught the vast majority of fractures with a sensitivity of 89.3%.
  • It maintained a specificity of 71.4%, keeping false alarms relatively low.
  • Performance remained stable across different sites, with an AUC degradation of less than 0.001.

This level of calibration is crucial. The model’s initial Brier score of 0.168 dropped to 0.151 after recalibration, with an expected calibration error of just 0.041. This means the probability scores the AI spits out actually match real-world risk, preventing the system from overwhelming clinics with false positives. Furthermore, the AI’s reasoning was highly transparent. The correlation between its quantitative scores and its written explanations exceeded 0.97 for three out of six clinical indicators.

The reality check

This is not a replacement for an MRI. A specificity of 71.4% means some healthy patients will still be flagged for scans they do not need. However, as a triage mechanism to prioritize who gets imaging first, this is incredibly practical. It shifts the diagnostic starting line from the radiology department to the patient’s living room.

The cohort size of 204 patients is still relatively small. Larger, more diverse trials are needed before this can be deployed in local clinics. But the ability of an LLM to spontaneously generate clinically coherent language to describe these movements suggests we are moving toward highly interpretable home diagnostics.

Read the full study in npj Digital Medicine.

Share on facebook
Facebook
Share on twitter
Twitter
Share on linkedin
LinkedIn
Share on whatsapp
WhatsApp

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.