← Back to AI Health Hub

AI predicts breast cancer survival without mimicking pathologists

New clinical data shows that AI models do not need to mimic human pathologists to accurately predict breast cancer outcomes.

New clinical data shows that AI models do not need to mimic human pathologists to accurately predict breast cancer outcomes.

Why do we insist that medical AI must copy human behavior to be valid? For years, developers have tuned algorithms to replicate the exact scoring patterns of pathologists. A new study from the PARTNER randomized trial upends this assumption, showing that while AI and humans disagree on the exact numbers, they arrive at the same clinical truth.

This disconnect is the real story.

For years, the field has chased human-like consensus to prove software works. This trial suggests the active clinical ingredient in AI diagnostics may be entirely different from human observation. It challenges the entire philosophy of “pathologist-in-the-loop” design.

The human-machine mismatch

Researchers tested two new pipeline updates, TRIGS and SAM-TIL, alongside older models like HoVerNet and muTILs. They compared these tools against gold-standard pathologist scores across multiple patient cohorts. The correlation between the AI models and human pathologists was surprisingly low, ranging from only 0.59 to 0.69 in 285 patients from the PARTNER trial. This is too low for the AI to act as a direct assistant or second opinion. The models are simply not interchangeable with human eyes.

Equal performance, different paths

Yet, when it came to predicting actual patient survival and treatment response, the AI matched the experts step for step. In the TransNEO cohort of 166 patients, both new models successfully predicted pathological complete response. TRIGS achieved an odds ratio of 1.95, while SAM-TIL reached 2.32. The predictive power of these models is remarkably consistent:

  • An odds ratio of 2.32 for SAM-TIL and 1.95 for TRIGS in predicting treatment response.
  • An overall survival hazard ratio of 0.80 for SAM-TIL and 0.79 for TRIGS in 277 TCGA patients.
  • A matching Area Under the Curve (AUC) of 0.60 to 0.64 for predicting treatment response across all tested methods and human experts.
  • An identical Integrated Brier score of approximately 0.10 for predicting event-free survival.

Stop copying, start validating

This parity in prediction, despite the low correlation, suggests that AI models see different, perhaps more complex, biological signals than humans do. This aligns with earlier findings in Discover Oncology, which highlighted the predictive power of deep-learning-based T lymphocyte quantification. Instead of forcing AI to follow manual scoring guidelines, the industry should validate these algorithms as independent diagnostic tools. We must stop treating human consensus as the only path to clinical utility.

The major limitation is that these models cannot yet support pathologists in their current workflow. Because they do not think like humans, they cannot easily explain their decisions to a human peer. They must be validated for autonomous use, which requires a massive shift in regulatory and clinical trust.

Read the full preprint in medRxiv.

This article is for informational purposes only and is not a substitute for professional medical advice, diagnosis or treatment.