Source-linked AI summary

The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

Siddharth Vohra, Manikandan Ravikiran

arXiv:2609.09048v1cs.CLcs.AIcs.CY

TL;DR

The paper asks whether apparent demographic bias in language-model decision audits depends on audit format, extending a reported rating-versus-ranking reversal beyond charitable aid. It tests hiring, lending, and triage with 40,726 requests to five models using controlled names, formats, and ranking constructions. The reversal does not generalize under the tested instrument: demographic contrasts do not survive correction, while recognition, ties, and candidate position reveal stronger audit-construction effects.

  • Problem

    The paper addresses whether conflicting demographic-bias findings reflect model behavior or the format and construction of the audit question.

  • Method

    The study sends 40,726 controlled hiring, lending, and triage requests to five models, varying names, rating or decision formats, and four-candidate ranking compositions.

  • Results

    None of 36 planned contrasts survives correction across the three domains, while models tie identical-content comparisons and show strong audit and position effects.

  • Takeaways & Limitations

    Under this tie-permitting instrument, audit verdicts reflect construction features such as recognition, wording, and candidate position more than measured demographic contrasts.

  • Takeaways & Limitations

    The conclusions use names as the only demographic signal, rely on few base profiles and five models, and exclude effects below the study's detection floors.

Abstract

from arXiv · show

Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary test fixed before collection. It does not. None of 36 planned contrasts survives correction. The rating advantage keeps its sign at roughly half the published size, and a precision extension bounds any hiring ranking penalty below the published effect, though the lending and triage ranking floors sit above that margin, so the exclusion is conclusive for hiring ranking and for rating in all three domains only. Planted disparities tracking their injected sizes and a directional replication on the original aid materials bound these nulls. The audit is livelier than the demographics: models recognize transparent audits nearly always, tie every identical-content comparison whether the varying detail is race or a hobby, and reward first-listed candidates as much as any demographic effect we measure. Audit verdicts reflect audit construction more than demographic bias.

1 Introduction

Prior audits disagree about demographic effects in consequential language-model decisions because they vary models, stimuli, cues, measures, and question format together. This study tests whether a format-driven reversal found in charitable aid generalizes to hiring, lending, and triage.

  • Legal-stakes applications include job scoring, credit recommendations, and triage assessments that can pass downstream without case-by-case human review.
  • Prior studies report both favoritism and penalties toward minority or female applicants, but their designs cannot isolate which audit choice contributes to the disagreement.
  • In charitable aid, the same models favored minority claimants when rating individually but penalized some when ranking side by side, motivating a cross-domain test of format effects.
  • The study applies name-based correspondence audits to hiring, lending, and triage, using identical-content applications that differ only in names signaling race and gender.
  • Across the planned findings, the reversal fails to appear, while audit recognition, ties, position effects, and wording artifacts emerge as prominent instrument effects.

2 Method

The study constructs controlled demographic and ranking audits across three decision domains and evaluates them with five models and preregistered contrasts. Its design varies presentation format, comparison structure, and audit transparency while preserving identical underlying profiles.

  • Stimuli: The study uses 12 demographic-neutral base profiles per domain and renders each under 40 validated name pairs, producing 480 single-profile stimuli per domain.
  • Factorial bundles: Ranking bundles cross demographic contrast with profile composition, comparing matched, irrelevant-attribute, disguised, and placebo conditions.
  • Formats: Rate requests a 1–5 score, Decide requests a binary outcome, and Rank requests a priority ordering with ties permitted and recorded.
  • Models and inference: Five models process 40,726 requests, with the primary hiring reversal test fixed before collection and mixed-effects regressions supplemented by rank-ordered-logit estimates.
  • Equivalence testing: Equivalence margins are fixed at 0.090 rating points and 0.067 rank positions, matching published effects used for exclusion tests.

3 Results

Across hiring, lending, and triage, the predicted rating-to-ranking reversal does not appear: no planned contrast survives correction, while rating effects retain a smaller positive sign and ranking penalties remain bounded only in hiring. Models also respond strongly to audit construction, tying identical-content comparisons and favoring first-listed candidates.

  • Primary test: +0.01 [-0.15, +0.17] for the format-by-group interaction provides no evidence of the predicted reversal.None of the 36 planned contrasts survives correction, and none of the twelve format-by-group interactions is distinguishable from zero.
  • Equivalence and sensitivity: +0.03 positions [-0.15, +0.20] on disguised Rank does not exclude the published -0.067 Asian disparity.A hiring extension narrows the pooled Rank contrast to +0.008 [-0.044, +0.060], bounding any hiring ranking penalty at 0.052 positions against the 0.067 margin; lending and triage floors remain above that margin.
  • Audit effects: 100% of demographic-matched bundles and 100% of hobby-or-neighborhood bundles receive ties, compared with 0% of disguised and placebo bundles.Models call the irrelevant-attribute cell a fairness or bias test 89% [84, 93] of the time, as often as the demographic cell.
  • Equivalence and sensitivity: 0.48 of an injected 0.5 disparity is recovered in the planted-control sensitivity reference, while every Figure 1 demographic interval covers or hugs zero.Planted disparities track their injected sizes, supporting interpretation of the nulls as bounded by measurable sensitivity rather than as proof of exact equality.
  • Audit effects: 0.094 standard deviations for first placement exceeds the largest demographic term of 0.049 on the same ranking outcome.The fitted difference is +0.045 [-0.167, +0.057], which covers zero, so the audit effect is comparable to or larger than any measured demographic effect rather than reliably larger.
  • Decision formats: -0.01 [-0.07, +0.06] and -0.02 [-0.08, +0.04] are the fitted race-by-wording interactions after applications are shared across groups.The apparent reversal in raw differences reflects different baseline applications, not wording once race and application are decoupled.
  • Decision formats: 56%, 56%, and 36% are the base rates for shortlisting, approval, and escalation, and every Decide contrast is null.Native function calls also change nothing: the submit_decision callback gap is +1.5pp [-1.2, +4.2].

4 Discussion and Conclusion

The study concludes that its audit fails to reproduce the aid-benchmark reversal across three regulated domains under a tie-permitting, quality-varied ranking design. Its evidence instead implicates benchmark materials and audit construction, while leaving the specific source of the original phenomenon unresolved.

  • Conclusion: Across hiring, lending, and triage, five models show no demographic contrast surviving correction under any tested format.The conclusion is limited to the study’s stated sensitivity and tie-permitting, quality-varied ranking construction.
  • Conclusion: The original aid materials reproduce both published effect directions, but neither interval clears zero.This points toward the materials without identifying whether domain semantics or stimulus construction carries the phenomenon.
  • Interpretation: A raw -0.32 second-candidate effect collapses to -0.01 after race and application are decoupled.The design therefore treats the apparent candidate effect as confounded by the construction of the compared applications, not as an established demographic effect.

Limitations

The study’s conclusions are bounded by its demographic signals, materials, models, detection floors, and controlled prompting. Several recognition and equivalence analyses also rely on limited or exploratory designs.

  • Scope boundaries: Names are the sole demographic signal, so results may not transfer to race or gender conveyed through dialect, explicit statements, or photographs.Name pairs also carry socioeconomic connotations that cannot be separated from race, and models’ perception of name cues is not measured.
  • Scope boundaries: The benchmark uses only 12 synthetic base profiles per domain and five models, limiting profile- and system-level generalization.Leave-one-profile-out and leave-one-name-out checks buffer but do not remove these scope limits.
  • Measurement limits: The wording arm covers one domain and two paraphrases per schedule, so its null bounds rather than establishes instruction sensitivity.Recognition rates are probe-dependent, and open-ended recognition uses a fixed keyword rule that was not independently validated.
  • Inference limits: Every null is bounded by the detection floor in Table 3, so effects below that floor remain unexcluded.The aid replication localizes the phenomenon to the original benchmark materials without identifying whether domain semantics or stimulus construction carries it.
  • Inference limits: Primary estimates use minimal prompts at temperature 0 with reasoning disabled where possible, while the reasoning-enabled rerun covers only two models and one domain.Temperature 0 is not determinism, and two endpoints inject uncontrolled sampling, so equivalence bounds lack run-level variance components.

Ethical Considerations

The study audits systems rather than people and uses public, fully crossed stimuli without human subjects. Its ethical framing also acknowledges that names carry associations beyond race and gender.

  • Study scope: The study involves no human subjects because it audits systems using stimuli from public releases.Groups are fully crossed with profile content, so no group is paired with weaker applications more often than another.
  • Study scope: Names are used as released, while the study acknowledges that they carry associations beyond race and gender, including perceived foreignness.This qualification matters when interpreting name-based demographic contrasts.

A Prompt Templates

The appendix fixes prompts across domains and operationalizes rating, binary decision, ranking, tool-call, and audit-recognition tasks with tightly constrained outputs. Prompt variants preserve the same task structure while changing wording or interface.

  • Prompt construction: Prompts are fixed before collection and shared across domains, with the decision verb instantiated as shortlist, approve, or escalate.The verbs correspond respectively to hiring, lending, and triage decisions.
  • Prompt construction: Rate requests require one integer on a 1–5 priority scale, with no explanation or reasoning.The appendix includes two paraphrased rating formulations using the same constrained output format.
  • Prompt construction: Decide requests require a binary integer indicating whether to shortlist, approve, or escalate the application.The output instruction permits only 1 for yes or 0 for no, without explanation or reasoning.
  • Prompt construction: Rank requests present four applications and require four comma-separated priority integers in presentation order, allowing equal ranks.Two paraphrases retain the same four-candidate ranking structure and tie policy.
  • Interface variants: Tool-call variants replace textual output instructions with native function schemas for one decision integer or four ranking integers.The Decide and Rank tasks use provider-native function-calling interfaces.
  • Recognition probe: Purpose probes show models the verbatim request and ask them to classify its likely purpose among processing, decision-quality, fairness, data-generation, or uncertainty options.The model returns only the option letter.

B Run Accounting and Configuration

The run accounting combines multiple task types within each model-domain cell and extends the primary runs with aid replication and targeted expansions. Configuration is largely standardized, with two provider-specific reasoning deviations.

  • Run accounting: Each model-domain cell contains 480 Rate, 480 Decide, 240 tool-call Decide, 600 Rank, 150 disguised Rank, and 48 probe requests.This totals 1,998 requests per model per domain and 5,994 per model.
  • Run accounting: Across five models, the primary study includes 29,970 requests, supplemented by 2,500 aid-replication requests and 8,256 extension requests.The extensions include hiring paraphrases, disguised hiring ranking, recognition probes, and further targeted runs.
  • Configuration: Configuration parity fixes temperature at 0, top-p at 1, and shared prompts, schemas, grammars, and pinned serving backends.Two enumerated provider deviations prevent complete symmetry across systems.
  • Configuration: Gemini 3.1 Pro cannot disable reasoning and therefore runs at vendor-minimum effort, averaging 327 thinking tokens per call.All other systems run with reasoning off, verified per call.
  • Configuration: Kimi K3’s pinned backend injects thinking tokens whenever a tool call is forced.This is the second provider-specific deviation from the shared configuration.

C Robustness

Robustness checks show that the null findings persist across specifications, clustering choices, wording controls, and name/profile exclusions. Precision extensions bound the hiring ranking penalty below the published margin, while position effects remain substantial and audit recognition complicates interpretation of transparent comparisons.

  • Wording robustness: -0.01 [-0.07, +0.06] and -0.02 [-0.08, +0.04] replace apparent wording reversals when race and application are analyzed on shared applications.A re-dealt arm with identical group quotas gives +0.01 [-0.05, +0.07] and -0.01 [-0.07, +0.05].
  • Estimator robustness: -87 to 114% changes in Rank interval widths from bundle-clustered standard errors do not change any conclusion.Narrowing occurs where ties leave little within-bundle variance for the naive estimator.
  • Interaction tests: None of the 12 format-by-group interactions survives Benjamini-Hochberg correction, including the preregistered hiring Black cell at +0.01 [-0.15, +0.17].The shared specification agrees closely with the preregistered estimate.
  • Position and ties: 0.11 positions [0.03, 0.18] is the estimated first-placement effect in disguised bundles, while the irrelevant-matched cell is fully tied and has no defined position effect.Position fixed effects are included in every Rank model.
  • Precision extension: 12,000 additional disguised hiring Rank slots narrow the pooled estimate to +0.008 [-0.044, +0.060] and bound any hiring Rank penalty at 0.052 positions against the 0.067 margin.The equivalence test returns p = .013.
  • Recognition probes: 100% [98, 100] and 74% [68, 80] false positives in unsampled probe cells show that forced-choice recognition can track the offered option rather than audit recognition.In open-ended probes, models identify 89% of transparent bundles and 89% of irrelevant-attribute bundles as fairness or bias tests.
  • Sampling robustness: 0.004 is the name-pair random-effect standard deviation, and leave-one-name-out or leave-one-profile-out changes no focal contrast by more than 0.02.No removal changes correction status.
Loading 2609.09048v1…