🧑🏼‍💻 Research - August 25, 2026

Large Language Models Generate Stigmatizing Language During Reasoning Over Real-World Clinical Data

🌟 Stay Updated!
Join AI Health Hub to receive the latest insights in health and AI.

AI models generate biased medical notes

Large language models trained to think through clinical decisions are actively generating and amplifying biased language that could compromise patient care.

We expect advanced AI models to filter out human prejudice, not mirror it. Yet, as health systems rush to adopt automated clinical documentation, they are introducing a silent vector for bias. This is not just a cosmetic issue. Stigmatizing language in electronic health records directly leads to worse clinical decisions and poorer patient outcomes.

The bias in the machine

A new study evaluated **107** large language models across **35** real-world clinical tasks, analyzing **3,745** model-task pairs. Researchers used a psychiatrist-validated natural language processing system to scan the models’ internal reasoning steps for judgmental or stereotyping terms.

The results shatter the assumption that “reasoning” models are more objective. Stigma rates ranged from **0% to 33.33%**, and a staggering **84.06%** of the model-task pairs contained stigmatizing terms.

Reasoning models perform worse

The data reveals a troubling paradox. Models designed for deeper reasoning actually generated more bias than simpler models.

  • Open-source models had higher stigma rates than proprietary models (**1.97% vs. 1.60%**).
  • Reasoning models showed significantly more stigma than non-reasoning models (**2.35% vs. 1.70%**).
  • Medical-specific models did not perform significantly better than general-purpose models (**1.80% vs. 2.00%**).
  • Worryingly, **19.76%** of model-task pairs actively amplified the bias present in the original patient notes.

This amplification is particularly dangerous. It suggests that AI does not just copy human flaws, it multiplies them. This aligns with recent findings on the challenges of AI-generated stigmatizing language regarding substance use disorders, showing how deeply these biases are baked into neural networks.

Stigma hurts accuracy

There is a direct cost to this bias. The study found a negative correlation of **-0.304** between stigma rates and task accuracy. When a model uses biased language, its clinical reasoning is objectively worse.

Fortunately, this is a solvable engineering problem. Applying targeted prompt engineering reduced model stigma rates by up to **91.91%** without degrading clinical performance. This proves that guardrails do not have to compromise utility.

However, we must acknowledge the limitations of this research. The study relied on automated detection on a preprint dataset, and real-world clinical environments are far more unpredictable. As researchers noted in previous work on efficient detection of stigmatizing language in electronic health records, identifying these subtle biases requires continuous, active validation. Health systems cannot treat AI safety as an afterthought. If we automate clinical workflows without strict destigmatizing protocols, we will institutionalize the very biases we have spent decades trying to eradicate.

This study was originally published in medRxiv.

Share on facebook
Facebook
Share on twitter
Twitter
Share on linkedin
LinkedIn
Share on whatsapp
WhatsApp

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.