A new foundation model uses baseline CT scans to flag which lung cancer patients are likely to develop life-threatening lung inflammation from immunotherapy.
Immunotherapy can save a lung cancer patient’s life, but it can also trigger a lethal side effect: pneumonitis. Doctors currently have no reliable way to predict who will suffer this immune-system backlash before treatment begins. This leaves clinicians guessing, balancing the promise of tumor shrinkage against the threat of sudden, severe lung damage.
This study complicates the current reliance on clinical risk scores. By extracting invisible patterns from standard pre-treatment CT scans, the CIPHER model proves that the risk of immunotherapy-induced pneumonitis (ICI-P) is already written into the lung’s baseline anatomy. It suggests that adverse events are not random immunological bad luck, but predictable structural vulnerabilities. This shifts the paradigm from reactive monitoring to proactive patient selection.
How CIPHER was built
Researchers pretrained CIPHER on 590,284 CT slices from 2,500 non-small cell lung cancer (NSCLC) patients to map lung tissue. They then adapted it to an internal cohort of 347 patients, where 33 developed adjudicated ICI-P. The model was fine-tuned using 254 non-ICI-P patients, leaving a held-out validation set of 93 patients (33 ICI-P cases and 60 controls) to test its accuracy.
Outperforming standard clinical tools
The results expose the weakness of traditional clinical forecasting. In head-to-head testing, CIPHER achieved an AUC of 0.83, completely outclassing the clinical model’s dismal AUC of 0.58. It also beat the radiomics model (0.77) and the ensemble model (0.78). This aligns with broader trends in oncology, where deep learning is outpacing traditional hand-crafted imaging features, as discussed in Artificial Intelligence and Radiomics.
To prove the model was not just memorizing local data, researchers tested it on an external cohort of 116 patients from Johns Hopkins, containing 20 ICI-P cases and 96 controls. CIPHER maintained its AUC of 0.83 and reached a balanced accuracy of 81.7%. Crucially, it showed a specificity of 83.3% compared to just 45.8% for the radiomics model, meaning it will not trigger constant false alarms in a busy clinic.
- Sensitivity of 80.0% in the external validation group.
- Correctly identified 80 of 96 non-ICI-P cases.
- Correctly flagged 16 of 20 active ICI-P cases.
- Significantly outperformed radiomics in external testing with a DeLong p-value of 0.0318.
The path to clinics
Despite these strong numbers, the study has limitations. It is a retrospective analysis, meaning the model must still prove its worth in prospective, real-time clinical trials. Furthermore, we must understand why the model makes these decisions. As explored in Explainable deep learning approaches, clinical trust requires transparency, not just high accuracy. If prospective trials validate these findings, CIPHER could allow oncologists to confidently prescribe aggressive immunotherapies to low-risk patients while routing high-risk individuals to safer alternative regimens.
Read the full study in medRxiv.
