A new AI analysis reveals that nearly half of all pregnancy records contain stigmatizing language, with Black and less-educated patients bearing the brunt of clinical bias.
When a pregnant patient walks into a clinic, the words written in their chart can quietly dictate the quality of care they receive. If a clinician labels a patient as “resistant” or “combative,” that judgment follows them to every future appointment. A new study using AI to scan hundreds of thousands of medical notes shows just how deeply this bias is baked into the system.
This is not just a matter of bedside manner. It is a systemic data integrity crisis. Electronic health records are treated as objective clinical truth, but they actually act as mirrors for human prejudice.
The scale of clinical bias
Researchers used a keyword-guided BERT classifier to analyze 640,345 obstetric notes from 26,178 pregnancies at a single academic medical center. The scale of the bias they uncovered is staggering. The algorithm flagged stigmatizing language in 47% of all pregnancies in the cohort.
This high rate suggests that biased documentation is not an occasional slip of the pen. Instead, it is a routine feature of the modern obstetric record. The clinical narrative is quietly compromised before any treatment even begins.
Who gets labeled
The distribution of this biased language is highly unequal, tracking closely with existing social disparities. The data shows clear patterns in which patients are targeted by clinical staff.
- Black patients had significantly higher odds of having stigmatizing language in their files compared to Asian patients (aOR = 1.5) or White patients (aOR = 1.4).
- Patients with only a 12th-grade education were far more likely to experience documented stigma than those with a college degree (aOR = 1.5).
- Patients who experienced indicated preterm births (aOR = 1.5) or spontaneous preterm births (aOR = 1.2) had higher odds of negative labeling compared to those who delivered at term.
The algorithmic feedback loop
These findings complicate the industry’s rush to build predictive AI tools for maternal health. If nearly half of all pregnancy records are tainted with subjective, stigmatizing language, any clinical algorithm trained on this data will inherit those same biases. We risk building tools that predict poor outcomes based on how much a clinician disliked a patient, rather than the patient’s actual biology.
There are limitations to this data. The study analyzed records from a single academic medical center, so the exact rates of stigma may vary across different regions and hospital systems. Yet, the sheer volume of notes analyzed makes the core trend hard to ignore.
Health systems must stop treating clinical notes as neutral data points. Using automated natural language processing to audit these records is a necessary first step, but the real work lies in retraining clinicians. Until we clean up the human input, the digital output will remain compromised.
Read the study in medRxiv.
