← Back to AI Health Hub

AI Patient Simulators Train Doctors on Biased Stereotypes

Medical schools are using AI to train future doctors, but a new audit reveals these digital patients are quietly reinforcing dangerous demographic stereotypes.

Medical schools are using AI to train future doctors, but a new audit reveals these digital patients are quietly reinforcing dangerous demographic stereotypes.

If an AI patient always links sexual risk to sexual orientation, is it actually teaching clinical judgment, or just automating prejudice? Large language models are increasingly used as digital standardized patients to train medical students. But these models do not just generate medical facts. They generate assumptions.

This challenges the assumption that AI-driven training is a neutral upgrade. It suggests that without strict separation of scriptwriting and live role-play, we are training doctors to rely on the very biases we have spent decades trying to eliminate. This aligns with broader concerns about algorithmic bias in public health AI, where systemic inequities are easily coded into automated tools.

Where the bias hides

The AI did not just mimic clinical reality. It exaggerated it. Researchers ran an audit using an HIV pre-exposure prophylaxis (PrEP) screening scenario. They tested 6 demographic factors across a 216-cell factorial design, simulating 4,320 conversations with fixed audit questions.

The data reveals how deeply stereotypes are baked into the AI’s logic:

  • Anal sex appeared in 100% of gay men’s cases, 83% of bisexual men’s cases, but only 3% of heterosexual men’s cases.
  • The six demographic factors explained 32% of the variance in composite sexual risk.
  • An AI patient’s education level predicted their assigned socioeconomic status with an adjusted R2 of 0.59.
  • Demographics predicted 5 of the 9 improvised probe responses in composite role-plays, and 2 under control cases.

Why this matters is highly specific to clinical training. If an AI patient always presents textbook stereotypes, a trainee never learns to screen a low-risk-profile patient who might actually need PrEP. They learn to associate risk with identity rather than behavior.

The double-masked trap

The real danger is how these biases mask each other. The study analyzed two distinct pathways: the written case script and the live, improvised role-play.

For example, the model wrote female characters with higher alcohol use in the scripts. Yet, during the live role-play, it portrayed them as drinking less. The two biases canceled each other out in the final composite.

If educators only audit the final conversation, they will miss both errors entirely. This highlights the complexity of managing AI with agency in clinical education. We cannot treat AI as a single black box.

This study is limited to a single tracer condition and a specific set of demographic factors. It remains a preprint and needs validation across other medical scenarios. Still, the implication is clear. If we deploy these tools unchecked, we risk graduating clinicians trained on digital caricatures.

Read the full study in medRxiv.