AI can now scan messy hospital notes to group cancer symptoms and predict which patients face the highest risk of death or readmission.
For decades, oncology has relied on patient surveys to track symptom clusters. These surveys are clean but rare, capturing only a single moment in time. Meanwhile, the real story of a patient’s decline is buried in the daily, messy narrative of clinical notes.
This study challenges the necessity of structured surveys. By letting off-the-shelf AI scan unstructured discharge summaries, researchers bypassed patient questionnaires entirely. The implication is clear: the natural language of clinical notes contains highly structured, predictive signals that we are currently ignoring. It suggests that we do not need to interrupt patients with constant forms to understand their trajectory.
Mining the messy notes
To prove this, researchers used two lightweight language models, Gemini 3.5 Flash and Claude Haiku, to analyze 2,728 discharge notes from 1,507 colorectal cancer patients in the MIMIC-IV database. The AI pipeline extracted 46 distinct symptoms with a baseline Macro F1 accuracy score of 0.70. Instead of relying on rigid patient surveys, the models mapped how these symptoms co-occurred naturally in the text. This approach treats the clinical note not just as a record, but as a live dataset.
What the AI found
The models grouped the symptoms into three distinct, clinically logical clusters: Systemic, Gastrointestinal, and Disease-Specific. The consistency between the two different AI models was remarkably high, showing a cross-model Adjusted Rand Index of 0.727. This high agreement suggests that LLMs are not just guessing. They are identifying a stable, underlying clinical reality in the text.
Here is how those automated symptom clusters mapped to real-world patient survival and hospital returns:
- The Systemic Symptom Cluster was the strongest predictor of death, raising the odds of in-hospital mortality by 1.33 times and one-year mortality by 1.41 times.
- The Disease-Specific Cluster independently predicted whether a patient would be readmitted within 30 days, increasing those odds by 1.20 times.
- These risk associations remained significant even after the researchers adjusted for age, sex, and whether the cancer had spread.
The real-world friction
This study proves we do not need custom-built, expensive clinical AI to extract valuable prognostic insights. Off-the-shelf models can do it. However, we must be honest about the limitations. The study relies on retrospective data from a single database, and a F1 accuracy score of 0.70 means the AI still misses or misinterprets a portion of the clinical text. It is a proof of concept, not a finished diagnostic tool.
If health systems want to use this to flag high-risk patients in real time, they must accept some noise in the data. But waiting for perfect manual surveys is a luxury clinics cannot afford. Using AI to turn passive documentation into active risk-scoring is a trade-off we should start making today. It shifts the burden of data collection from the sick patient to the computer.
Read the full preprint study in medRxiv.



