Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews LLM Decision-Making
Abstract
Clinicians frequently use stigmatizing language (SL) in medical notes—expressing doubt, blame, or maligning—which is known to skew human clinical decisions. As large language models (LLMs) enter clinical workflows, we investigated whether they inherit these biases. We evaluated nine frontier LLMs across four stigmatized conditions (sickle cell disease, obesity, cirrhosis, fibromyalgia), comparing neutral clinical vignettes against versions injected with varying doses of SL while holding all objective clinical data constant. All nine models shifted toward less aggressive management when exposed to SL, with a single stigmatizing sentence sufficient to alter decisions and a clear dose–response relationship. Simulated clinician attitudes declined uniformly across all models and conditions. Strikingly, SL influenced outputs far more than patient demographics (age, gender, race). Prompt-based mitigation, including chain-of-thought and self-debiasing, provided only partial relief. These findings reveal that LLMs propagate implicit linguistic bias, threatening health equity and demanding validated safeguards before clinical deployment.