← Back to AI Health Hub

AI swarms overestimate ICU mortality risk

Letting AI models revise their own clinical decisions makes them more pessimistic and biased, not more accurate.

Letting AI models revise their own clinical decisions makes them more pessimistic and biased, not more accurate.

We expect medical AI to get smarter when it takes a second look. But a new evaluation of adaptive Large Language Model (LLM) swarms reveals a troubling trend.

Instead, they double down on bad news.

When AI agents debate ICU mortality risks, they do not find hidden nuances. They ignore counterevidence to justify their initial hunches. This challenges the belief that multi-agent “reasoning” loops make clinical AI safer.

The bias toward death

Researchers analyzed 1,607 eICU encounters (representing 1,500 stays) and 1,607 ICU-2012 encounters to see how adaptive LLM swarms revise their decisions. The system allowed specialist AI agents to review and update risk scores. Revision happened often, occurring in 84.32% of eICU cases and 56.44% of ICU-2012 cases.

But these revisions did not make the predictions more accurate. There was no clear improvement in standard performance metrics like AUROC or AUPRC. Instead, the AI simply became more pessimistic.

Mean risk scores increased by +0.0810 in the eICU group and +0.0539 in the ICU-2012 group. Remarkably, 757 of 761 threshold crossings moved toward predicting mortality.

More evidence, less balance

The study compared these adaptive outputs to a simpler “fixed-voting” system. While the revising swarm looked more thorough on the surface, its internal logic was highly defensive.

The final revised outputs had 1.26 and 0.55 more supporting-evidence items than the fixed-voting model. However, they actively ignored conflicting data. Counterevidence acknowledgment dropped by 21.59 and 7.47 percentage points.

At the same time, unsupported-claim flags rose by 1.43 and 0.68 points.

  • Revisions occurred in up to 84.32% of patient encounters but failed to improve overall predictive accuracy.
  • Almost all risk-score changes (757 out of 761) shifted toward predicting death.
  • The revision process nearly doubled the computational burden, requiring up to 1.97 times the runtime of simpler models.

The cost of overthinking

This behavior mimics a well-known human flaw: confirmation bias. Instead of weighing all facts objectively, the revising AI gathered more points to support its thesis while ignoring warnings.

This extra “thinking” also came with a heavy computational cost, taking 1.97 and 1.52 times the runtime of the simpler method.

For clinicians, this is a warning. More complex AI architectures do not guarantee better clinical insights. If an AI system spends extra time and energy to become more biased, we are better off sticking to simpler, faster models.

This study was limited by its retrospective design and did not evaluate real-time clinical decision-making or actual patient outcomes.

Read the full study on medRxiv.