Source-linked AI summary

Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice

Syeda Anshrah Gillani, Mirza Samad Ahmed Baig

arXiv:2608.14399v1cs.CYcs.AIcs.CL

TL;DR

LLM assistants increasingly mediate physician choice, but it remains unclear whether demographic signals affect recommendations when reputational attributes are held fixed. This prespecified randomized audit finds that ratings and price dominate, while small demographic tilts favoring female- and minority-signaled names remain absent from model explanations.

  • Problem

    The paper asks whether physician names signaling gender or ethnicity affect LLM recommendations when reputational attributes are held fixed.

  • Method

    The authors use a randomized choice-based conjoint audit of LLMs selecting among synthetic physician cards with independently randomized attributes.

  • Results

    Female- and minority-signaled names gain one to three percentage points, while these demographic effects appear in fewer than 0.03% of stated reasons.

  • Takeaways & Limitations

    Demographic tilts are not observable from model self-report but are measurable cheaply and repeatably with a frozen audit design.

  • Takeaways & Limitations

    The design holds specialty, board certification, and distance constant, so effects of those attributes are outside its scope.

Abstract

from arXiv · show

Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models' own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.

1 Introduction

This paper treats LLM physician recommendations as an AI-infomediary problem and uses a prespecified randomized conjoint audit to identify which reputation, cost, demographic, and presentation signals move choices. Across seven models, reputation dominates, demographic parity fails in an unexpected direction, and stated explanations largely miss the revealed effects.

  • Audit design: The audit randomizes physician-card attributes independently across synthetic profiles, identifying each signal’s average marginal component effect and converting effects into fee-equivalent dollars per visit.Gender and ethnicity are signaled through physician names using correspondence-audit methodology, while display order is also randomized.
  • Audit design: The frozen instrument makes findings falsifiable and repeatable by allowing future models to be audited against identical stimuli, including parity, abstention, and explanation tests.The study spans seven models, three patient personas, nine prompt paraphrases, and nine prespecified experimental arms.
  • Main findings: Rating carries 37.7% of total attribute importance and fee 24.1%, with a 3.9→4.7 rating increase worth $157 per visit in fee-equivalent terms.Across seven models and 40,068 scored responses, reputation is the dominant basis of physician recommendation.
  • Main findings: Female-signaled and minority-signaled names gain one to three percentage points, worth $7–$14 per visit, while no gender×ethnicity cell meets the prespecified equivalence bound.Demographic parity therefore fails, but not in the direction predicted by human audit studies.
  • Main findings: Gender and ethnicity appear in fewer than 0.03% of stated reasons, and models abstain in only 0.39% of trials, leaving demographic and position effects invisible in self-reported rationales.The paper therefore compares revealed recommendation weights with the models’ explanations rather than treating explanations as sufficient evidence of behavior.

2 Background and hypotheses

The study tests whether LLM physician recommendations reproduce established reputation and access preferences while allowing demographic name signals to affect choices. It prespecifies parity, intersectional, context-sensitivity, position, abstention, and explanation tests rather than assuming a directional demographic effect.

  • Reputation and choice attributes: Prior conjoint evidence predicts that higher ratings and service ratings increase physician choice, while fees, access, and credentials matter according to patient circumstances.The study’s personas mirror heterogeneity such as uninsured patients weighing fees differently from chronically ill patients.
  • Reputation and choice attributes: The prespecified H1–H8 hypotheses predict positive recommendation effects for ratings, review volume and recency, responsiveness, affiliation, telehealth, experience, and lower fees.The expected ordering is valence first, price next, followed by smaller positive secondary signals.
  • Demographic signals: Abstention and explicit acknowledgment of demographic attributes are treated as outcomes because declining to choose may be more defensible than silently using names.The design follows evidence that names can shift decisions even when qualifications are identical.
  • Demographic signals: The study asks whether gender- or ethnicity-signaling names change recommendations among otherwise identical physicians, distinguishing inherited discrimination, neutrality, and overcorrection.Parity is tested jointly with equivalence bounds of ±1.5 percentage points, alongside gender×ethnicity cell effects and an additivity test.
  • Context and auditability: Additional hypotheses test whether content-free display position affects recommendations, whether fee sensitivity varies by persona, and whether review volume moderates the rating effect.The study also examines whether stated reasons track revealed choice weights, especially when demographic attributes influence choices.

3 Methods

The audit uses a frozen, randomized conjoint design in which models choose among five synthetic family physicians whose reputational and demographic signals vary independently while clinical comparators remain constant. Identical stimuli, prespecified secondary arms, and pooled AMCE-based inference support within-trial model comparisons and causal attribution of choice differences to signals.

  • Choice design: Each trial presented one model, patient persona, prompt template, and five synthetic family-physician cards, holding specialty, certification, distance, and new-patient status constant.Eight reputational attributes varied independently and uniformly, including ratings from 3.9 to 4.7, fees from $90 to $190, reviews, recency, feedback response, affiliation, telehealth, and experience.
  • Demographic signals: Gender and ethnicity were randomized per card and signaled through names drawn from correspondence-audit conventions, with cell-level analysis, placebo tests, and unique surnames within each choice set.The demographic cells included female or male signals and White, Black, Hispanic, East Asian, or South Asian signals.
  • Experimental arms: The design comprised 3,024 fixed main-arm choice sets crossing three personas with nine paraphrased prompts, and every audited model received identical stimuli.Eight secondary arms altered conditions such as temperature, web-snippet formatting, ranking, field order, repetition, position controls, persona presence, and card-line order.
  • Pre-registration and inference: The design matrix and analysis plan were generated, hashed, and frozen before confirmatory collection, with estimands defined as attribute-level AMCEs on recommendation probability.Primary inference uses linear-probability regressions with slot dummies and choice-set-clustered errors, estimated per model and pooled with model fixed effects; conditional logit is secondary.
  • Audited models and exclusions: The audited panel included six local open-weight instruction-tuned models and gpt-4o-mini under the identical protocol, while deepseek-r1:7b was excluded by a prespecified pilot gate.Claude-family pilot records used a non-reproducible interactive harness and were quarantined from confirmatory analysis, serving only descriptive and future protocol-comparison purposes.

4 Results · 4.1 Data quality and manipulation checks · 4.2 What moves the recommendation: AMCEs

Across 40,068 scored responses, the audit produced high-quality, repeatable recommendations and showed that reputation signals—especially ratings and fees—dominate physician choice. Most prespecified attribute hypotheses were supported, while hospital affiliation was not.

  • 4.1 Data quality and manipulation checks: The confirmatory panel included seven models—six open-weight instruction-tuned models and gpt-4o-mini—and yielded 40,068 scored responses across nine arms.The open-weight models were llama3.2:3b, qwen2.5:3b, phi3:mini, mistral:7b-instruct, gemma3:4b, and llama3.1:8b.
  • 4.1 Data quality and manipulation checks: Parse failures were below 2% for nearly all model–arm cells, but mistral reached 32.0% on retest and 14.1% on top-3 ranking, while deepseek-r1:7b failed the pilot gate entirely.gpt-4o-mini had 0% parse failures on every arm; deepseek-r1:7b had 100% parse failures.
  • 4.1 Data quality and manipulation checks: Test–retest consistency was high, with modal choices repeated in 73% of mistral repetitions and 99% of gemma3 repetitions.These percentages summarize five-fold repetitions.
  • 4.2 What moves the recommendation: AMCEs: A 0.8-point rating increase raises choice probability by 31.4 percentage points, while a $100 fee increase lowers it by 20.0 percentage points.These pooled effects are measured against a 20% baseline; the rating estimate has a 95% CI of 30.5–32.3 and the fee estimate a CI of 19.1–21.0.
  • 4.2 What moves the recommendation: AMCEs: Review volume, recency, telehealth availability, feedback responsiveness, and experience all increased choice in the expected direction.The reported effects were +8.4 pp for 400 versus 12 reviews, +3.8 pp for review recency, +2.8 pp for telehealth, and +2.1 pp each for feedback responsiveness and 28 versus 8 years in practice.
  • 4.2 What moves the recommendation: AMCEs: Rating accounts for 37.7% of total attribute-range importance, followed by price at 24.1%, review volume at 10.1%, and list position at 8.2%.Confirmatory hypotheses H1–H4 and H6–H8 were all supported after Holm correction.
  • 4.2 What moves the recommendation: AMCEs: Hospital-system affiliation had no detectable effect (+0.3 pp, P = .35), leaving prespecified reputation hypothesis H5 unsupported.

4.3 Name-signaled gender and ethnicity

Models reject demographic parity, showing a systematic pro-female and pro-minority choice tilt, although name-specific familiarity entangles these audit-level signals. They almost never disclose that demographics moved their choices.

  • Name-signaled gender and ethnicity: 2.5 pp more often chose female-signaled names, while Hispanic-, South-Asian–, and Black-signaled names gained 2.8, 2.9, and 1.3 pp versus White-signaled names.East-Asian–signaled names were indistinguishable from White-signaled names (+0.3 pp, P = .65).
  • Name-signaled gender and ethnicity: Choice shares differed across name exemplars within gender×ethnicity cells, so demographic estimates capture perceived-category signals entangled with name-specific familiarity.The prespecified name-exemplar placebo was violated (LR = 370.0, df = 30, P < .001), making the estimates audit-level signals rather than precise category parameters.
  • Gender×ethnicity combinations: Female-Hispanic cards gained 5.7 pp, female–South-Asian cards 5.1 pp, and female-Black cards 3.9 pp relative to White-male cards.The interaction was not jointly significant after correction, but no gender×ethnicity cell met the prespecified equivalence criterion.
  • Auditability: 0.39% of responses abstained, while models spontaneously flagged demographic attributes in only 0.01%, despite names moving their choices.The models that let names move their choices essentially never said so.

4.4 Position effects

Physician-card position produces economically meaningful, content-free choice effects: slots 4 and 5 underperform slot 1, while letter-label decomposition finds small isolated components and a residual anti-“E” preference.

  • Position effects: Cards in slot 5 lose 6.2 pp and slot 4 loses 3.3 pp relative to slot 1, holding all content constant.The slot-5 estimate has a 5.0–7.3 confidence interval; the comparison is reported in Figure 5.
  • Position effects: Both slot and letter components are small in isolation, with partial R2 = .005 for slot given letter and .008 for letter given slot.The decomposition uses gpt-4o-mini token logprobs in the letter-labeled arm.
  • Position effects: A residual anti-“E” letter preference remains at −7.3 pp (P = .03).This residual estimate is reported in Table 8.

4.5 Fee-equivalents · 4.6 Persona and interaction effects

Fee-equivalent estimates show that demographic and interface signals carry substantial monetary value, while persona and rating-volume interactions condition how fees and reputation affect physician choice. Price sensitivity is strongest for uninsured personas, and high review volume amplifies the premium for highly rated physicians.

  • 4.5 Fee-equivalents: $157 per visit: the 3.9→4.7 rating increase is the largest fee-equivalent signal, with a 95% CI of $147–$167.The conversion uses the monotonic fee gradient required by the prespecified gate.
  • 4.5 Fee-equivalents: $19, $14, and $10 per visit: recency, telehealth availability, and responding to feedback receive these fee-equivalent values, respectively.These estimates are derived through the same fee-gradient conversion.
  • 4.5 Fee-equivalents: $14 per visit: Hispanic- and South-Asian–signaled names carry this fee-equivalent advantage, while female-signaled names carry $12.50 and first-listed cards $11.The first-listed effect is described as a pure interface artifact priced like a clinical credential.
  • 4.5 Fee-equivalents: Demographic advantages are evaluated against White-male cards using 95% confidence intervals and a prespecified ±1.5 pp equivalence band.Figure 4 reports the gender×ethnicity cells relative to the White-male reference.
  • 4.6 Persona and interaction effects: The uninsured persona’s price penalty significantly exceeds those of insured personas (H11, PHolm < .001).The persona-dependent fee sensitivity is shown in Figure 6.
  • 4.6 Persona and interaction effects: High review volume amplifies the high-rating premium (H12, PHolm < .001), and the sign replicates in the conditional logit.Probability-scale interactions require caution under a logit-additive process, as noted in Methods.

4.7 Stated versus revealed importance

Stated reasons track reputation but largely conceal demographic and position effects. Although rating and price are frequently mentioned, model self-report would have detected none of the measured demographic or position effects.

  • Stated versus revealed importance: Rating appears in 57–91% of stated reasons and price in 42–56%, roughly matching their revealed importance.The analysis coded 21,143 reasons across models.
  • Stated versus revealed importance: A self-report-based explanation mandate would have detected none of the demographic or position effects measured here.The conclusion applies to all demographic and position effects examined in this section.

4.8 Robustness

The reputation hierarchy remains stable across prespecified perturbations, with no sign changes and only small shifts in headline AMCEs. Removing personas increases rating sensitivity and reduces the fee penalty, consistent with budget information carried by personas.

  • Perturbation stability: Across all prespecified perturbations, reputation effects shift by at most a few points without changing sign, and temperature-0 decoding preserves their ordering.Robustness holds under grounded web-snippet formatting, reversed card layout, field-order changes, and the plausibility-filtered subsample.
  • Persona robustness: +8.4 pp on the 4.7 rating step and −9.7 pp for the fee penalty occur when personas are removed.The changes are consistent with personas carrying real budget information.
  • Missingness: Parse failures are unrelated to card content, according to the missingness model.
  • Exceptions: The gate-excluded deepseek-r1 and the mistral retest arm are the only material robustness exceptions.

4.9 Exploratory: frontier-model session pilot

An exploratory pilot of four Claude-family models reproduced the audit’s reputation and fee ordering, while revealing directional demographic patterns that require protocol-compliant confirmation. The pilot was descriptive, quarantined from confirmatory analysis, and its archived forecasts remained ungraded.

  • Pilot design: Four Claude-family models completed 200/200 pilot calls without parse failures, using the first 50 choice sets of the frozen design.Responses were collected interactively in an agentic coding harness with default decoding, not through the audit-protocol Messages API.
  • Limitations: Archived pilot-based full-run forecasts could not be graded because no protocol-compliant API run of a forecast-target model was available at analysis freeze.The forecasts remain timestamped and available for scoring when such a run occurs.
  • Exploratory results: 39–43% of cards rated 4.7 were chosen versus a 20% null, while 3.9-rated cards were chosen 0–1%; $90 fees yielded 35–37% choice versus 5–8% for $190.Higher review volume and recency also shifted choice in the expected directions across all four models.
  • Exploratory results: Female-signaled names were chosen 22–24% of the time versus 16–19% for male-signaled names, directionally matching the confirmatory panel’s pro-female tilt.At 50 choice sets per model, this pattern was directional and not hypothesis-tested.
  • Exploratory results: East-Asian–signaled names were chosen 31–36% versus 12–19% for White-signaled and 9–13% for Black-signaled names, unlike the confirmatory panel.These directional readings require confirmatory scrutiny and were not hypothesis-tested.

5 Discussion

The audit finds that AI physician recommendations are driven chiefly by ratings and price, with additional opaque position and demographic tilts that self-reported explanations do not reveal. Because these effects vary across models and versions, frozen-design behavioral audits should recur at each release rather than rely on explanation mandates.

  • Core findings: 37.7% of attribute importance comes from rating and 24.1% from fee, with rating weighted at $157 per visit and signals steeper than human benchmarks.The rating and fee weights are directionally consistent with human physician-choice conjoint estimates but exceed them.
  • Core findings: Female-, Hispanic-, South-Asian–, and Black-signaled names gain one to three percentage points over male White-signaled names, worth $7–$14 per visit.No gender×ethnicity cell meets the prespecified ±1.5 pp equivalence bound.
  • Core findings: A content-free first-listed position is worth $11 per visit, while demographic and position effects never appear in the models’ stated reasons.These effects are therefore not observable through model self-report.
  • Audit implications: The frozen randomized instrument measures causal signal weights cheaply and repeatably, supporting recurring behavioral audits because model updates can silently reweight signals.Audits are proposed as per-release monitoring rather than one-shot evaluation, with each new model release auditable in a day.
  • Limitations: External validity is limited by synthetic text-only profiles, name-based perceived-category signals, six small open-weight models plus gpt-4o-mini, and snapshot-specific versions.The exploratory frontier pilot used default decoding and n = 50 choice sets per model, and ±1.5 pp equivalence bounds may miss smaller systematic effects.
Loading 2608.14399v1…