A new AI-driven analysis of 50,000 patients reveals that severe cancer treatment side effects are not random mishaps but predictable genetic events.
Oncology has a long-standing blind spot. Doctors are highly skilled at choosing drugs to destroy tumors, but they remain remarkably poor at predicting which patients those same drugs will severely harm. For decades, the medical community has treated severe toxicity as an unavoidable, unpredictable tax on survival.
A new preprint from Memorial Sloan Kettering directly challenges this passive stance. By using a large language model to parse unstructured clinical records, researchers built MSK-Tox, a massive dataset tracking side effects across more than 50,000 patients. The resulting analysis shows that severe reactions are often written into a patient’s DNA long before they receive their first infusion.
Two paths to toxicity
The study focused on six representative adverse events: pneumonitis, adrenal insufficiency, liver toxicity, colitis, hyperthyroidism, and hypothyroidism. The researchers discovered that genetic risk is not a monolith but falls into two distinct biological categories.
First, some patients possess an organ-intrinsic vulnerability. For example, a genetic variant near the FOXE1 gene predisposes patients to hypothyroidism across various types of systemic therapies. Second, other risks are immune-mediated and only surface during specific treatments. The researchers found that the HLA-DRB1*15 allele is a major driver of adrenal insufficiency, but only in patients treated with immune checkpoint inhibitors.
This is the critical analytical takeaway. The same HLA-DRB1*15 allele is already known to predispose individuals to multiple sclerosis. This means checkpoint inhibitors do not create a new disease out of thin air. Instead, the therapy unmasks a latent, pre-existing autoimmune predisposition.
Shifting to precision safety
This finding changes how we must evaluate drug safety. For years, the pharmaceutical industry has accepted side effects as collateral damage. These results suggest we should screen patients for genetic risks before prescribing heavy immunotherapies, a concept central to the growing field of immunopharmacogenomics.
By combining pretreatment clinical features with genetic data, doctors could theoretically forecast a patient’s specific hazard profile. This moves oncology closer to a true personalized medicine approach, where safety is calculated as precisely as efficacy.
The limits of data
- Analyzed clinical and genetic data from over 50,000 patients.
- Mapped six major treatment toxicities including liver damage and colitis.
- Identified two distinct genetic modes of susceptibility: organ-intrinsic and immune-mediated.
- Linked the HLA-DRB1*15 allele to a high risk of drug-induced adrenal insufficiency.
However, we must view these findings with healthy skepticism. Because the data relies on retrospective electronic health records parsed by an AI model, there is always a risk of misclassification. Furthermore, the cohort is heavily drawn from a single, highly specialized cancer center, which may not represent the broader, more diverse global population.
Even with these limitations, the study proves that patient harm is not a mystery. It is a biological variable we can measure, predict, and eventually avoid.
Read the full preprint on medRxiv.



