Source-linked AI summary

Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts

Diego Mardian, Frank Liu

arXiv:2608.15254v1cs.AIcs.CL

TL;DR

Equity-aware prompts are increasingly recommended for clinical language models, but it is unclear whether they can make models add unstated patient demographics. Across 47 models and four medical benchmarks, this study finds that a single DEI prompt raises demographic injection from 0.7% to 33.1% in every model tested.

  • Problem

    The paper asks whether equity-aware prompting causes clinical language models to add unstated demographic attributes, potentially misrepresenting patients and changing recommendations.

  • Method

    The study evaluates 47 models on 376,000 responses across four medical benchmarks using matched prompt conditions and validated model-judge scoring.

  • Results

    A DEI prompt raises demographic injection from 0.7% to 33.1% in all 47 models, exceeding a length-matched neutral control by a median 18×.

  • Takeaways & Limitations

    DEI prompts can misrepresent patients, and a small subset of injected demographics changes the recommended answer toward the incorrect option.

  • Takeaways & Limitations

    Answer-altering rates are near-floors, gold labels are single-annotator, benchmarks are multiple-choice, and the proposed mechanism lacks activation-level analysis.

Abstract

from arXiv · show

Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes (race, socioeconomic status, sex) the question never stated, in effect rewriting who the patient is. We call this demographic injection. Across 47 models, four medical benchmarks, and 376,000 responses scored by a validated model-judge pipeline, a single DEI prompt raises the injection rate from 0.7% to 33.1% (47x) in all 47 of 47 models, attributable to the equity content rather than to added length (18x above a length-matched control; p=1.4x10^-14). Most added content is a general population statement that leaves the answer unchanged, but a smaller subset attaches an attribute to the specific patient or changes the selected option (0.25-2.4% of responses, 99.8% toward the incorrect option), where the invented demographic changes the answer the model recommends. Phrasing scales the effect from 14% to 56%. DEI prompts are just one example of a more general mechanism. Any instruction that nudges how a model reasons can make it add unrequested details, including details about the patient. Flagged outputs are treated as model errors under study, not clinical guidance.

1 Introduction

Equity-aware prompting can cause clinical language models to invent demographic details absent from a medical question, systematically altering what they write. The effect is mostly benign but can occasionally change the answer and may generalize to other framing directives.

  • Demographic injection: A one-sentence DEI prompt appended to a demographic-free medical question frequently causes models to add unstated race, age, social situation, or care-access details.This quantifies a consequence of increasingly recommended equity-aware prompting in clinical language models.
  • Study scope: Across 47 models, the study measures demographic injection and distinguishes mostly benign additions from occasional answer-altering injections using a higher-precision judge.The prompting practice systematically changes model output, while most changes do not alter the answer.
  • General mechanism: The authors argue that this effect may generalize beyond DEI prompts to any directive that frames how a model should reason.The broader concern is that framing instructions can systematically alter model-written content.

2 Method

The study compares four one-sentence addendum conditions across 47 models and four medical QA benchmarks, using model-judge scoring with conservative arbiter rescoring. It measures demographic injection and tests its sensitivity to DEI phrasing.

  • Experimental conditions: 47 models answer 500 items from four medical multiple-choice sets under baseline, nonsense, neutral-filler, and DEI addendum conditions.The addenda differ only in the appended one-sentence directive; nonsense controls for added instruction, while neutral filler controls for added length.
  • Scoring pipeline: Every response is labeled for injection, attribution, factual status, and answer influence by a deep judge, with candidate subsets conservatively re-scored by a higher-precision arbiter.The study pools all four benchmarks, and the subset rates are conservative near-floors.
  • Injection measurement: 0.7% to 33.1%: the DEI prompt raises the rate of added unstated patient demographics across all 47 models, far above nonsense and length-matched neutral controls.The rates are pooled over all four benchmarks.
  • Prompt phrasing: 14% to 56%: DEI-prompt phrasing changes injection frequency across the 47-model average.The dark portion represents cases where the model assumes an attribute about the specific patient.

3 Results

A single DEI prompt sharply and universally increases demographic injection across models, driven by equity content rather than added length. Most injections leave answers unchanged, but a small subset alters patient interpretation or answer selection, and phrasing substantially scales the effect.

  • Magnitude and mechanism: 0.7% to 33.1% (47×) in 47/47 models: a single DEI prompt increases unstated-demographic injection, while nonsense remains at baseline and DEI exceeds neutral control by median 18×.The effect is attributed to equity content rather than added text.
  • Clinical consequences: 91% of injections are general population statements, while 2.4% of all DEI responses change the answer, with 99.8% moving toward the incorrect option.Patient-directed attribution occurs in 0.71% of all DEI responses, and its intersection with answer-changing is 0.25%.
  • Prompt phrasing: 14% to 56%: phrasing the equity directive scales demographic injection, with verbose social-determinants framing maximizing patient-specific rewriting.Prompt design functions as both the mechanism and a control point.

4 Discussion

Demographic injection misrepresents patients and can reduce accuracy when an invented attribute changes the recommended answer. The discussion proposes instruction-following and learned associations as mechanisms, generalizes the risk beyond DEI prompts, and notes important study limitations.

  • Why this matters: Demographic injection attaches an unstated attribute to the patient, causing the model to reason about a different patient and sometimes select an incorrect answer.Across all 47 models, a small, consistent subset of these changes flips the selected option.
  • A likely mechanism: A proposed mechanism is that equity discourse co-occurs with demographic descriptors, increasing race- and SES-related continuations when DEI tokens appear.The effect scales with how explicitly demographics are requested and appears gated by instruction-following, consistent with treating the directive as a content request.
  • Why this matters: Only the DEI addendum in the example elicits demographic-group considerations and switches the diagnosis from correct option C to incorrect option D.The item concerns a 40-year-old woman with neck swelling, hyperthyroid symptoms, and normal ESR; the correct diagnosis is silent thyroiditis, while the switched diagnosis is Hashimoto’s thyroiditis.
  • Generalization: The risk likely generalizes beyond DEI: framing directives may elicit associated content unbidden, including insurance, liability, or demographic context that rewrites the patient and sometimes shifts the answer.Such injected framing may remain benign unless it becomes load-bearing for the answer.
  • Related work and limitations: The study separates benign from answer-altering injection but acknowledges near-floor answer-altering rates, single-annotator gold labels, multiple-choice benchmarks, and a mechanism awaiting activation-level analysis.The work frames the phenomenon as instruction-induced extrinsic hallucination rather than established demographic-bias measurement.
Loading 2608.15254v1…