Source-linked AI summary
Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models
William Guey, Pierrick Bougault, Wei Zhang, Vitor D. de Moura, José O. Gomes
TL;DR
The paper asks whether bias-audit instruments measure models comparably enough to rank them, and tests this with ten tools on shared model panels across workplace constructs. It finds reliable bias detection but no supported cross-tool model ranking, with disagreement tied to audit format and construct. The practical conclusion is that audits can characterize bias within an operationalization, but a single score should not be treated as a generally valid ranking.
Problem
Regulation and procurement increasingly use bias-audit scores, but evidence is lacking on whether different instruments agree enough to rank the same models.
Method
The study reimplements ten extrinsic tools in one harness and evaluates shared frontier and weaker-model panels across occupational gender, age, and socioeconomic-status bias.
Results
Cross-tool ranking agreement fails despite bias detection, while direction splits by format and the positive control recovers within-tool reliability but not cross-tool ranking.
Takeaways & Limitations
A single audit can detect bias and estimate direction within its stated operationalization, but it does not support ranking one model against another.
Takeaways & Limitations
The tools were reimplemented rather than run from their original codebases, so their numbers have not yet been validated against published scores on shared models.
Abstract
from arXiv · showhide
Emerging AI regulation mandates bias audits of high-risk systems, and audit scores are beginning to be used to rank models. Both uses assume different audit tools measure the same thing well enough to compare. We test that assumption directly, running ten extrinsic audit instruments over a shared panel of ten frontier models through one pooled inference gateway, first on occupational gender bias, then on age and socioeconomic status. Detection succeeds while ranking fails. Eight of ten tools detect bias with confidence intervals clear of zero; two widely cited direct-probe benchmarks are saturated because frontier models now answer neutrally. But cross-tool rank agreement is indistinguishable from chance (Kendall's W=0.07, p=0.83). A positive control with six deliberately weaker models separates two explanations: within-tool reliability recovers once the panel spans real capability gaps, yet cross-tool ranking never recovers, which points to the tools measuring different constructs rather than one construct noisily. Even the direction of bias splits by audit format: forced-choice decision tools mostly over-correct (toward women, and toward working-class candidates in 273 of 278 hiring decisions), while free generation and default coreference stay stereotype-congruent. The pattern replicates on socioeconomic status; an apparent ranking agreement on age dissolves under the paper's own tool-inclusion rules. The practical message: a single audit can detect bias and estimate its direction within its own operationalization, but no single audit supports ranking one model against another. All raw responses, code, and the analysis that recomputes every reported number from source are available at https://github.com/williamguey/bias-audit-agreement.
1 Introduction
The paper tests whether bias audits can support both bias detection and model ranking, holding models, construct, and inference access constant. It finds that tools detect bias but disagree on rankings and sometimes on direction.
- 1 Introduction: Regulatory and procurement uses increasingly rely on audit scores, assuming different instruments produce comparable numbers for ranking models.The motivating examples include EU and New York City employment-audit requirements and benchmark-based purchasing decisions.
- 1 Introduction: The literature contains dozens of instruments built around incompatible bias concepts, but their agreement on model rankings has not been tested on shared models.Existing tools span question answering, surveys, generation, coreference, and demographic-swap decisions, typically validated on separate model sets.
- 1 Introduction: Ten extrinsic instruments are compared on the same frontier-model panel, construct, and pooled gateway so tool identity is the planned source of variation.The study tests both reliability and cross-tool agreement rather than treating disagreement among unreliable instruments as informative.
- 1 Introduction: Tools agree that bias is present, yet their model rankings are statistically indistinguishable from random and their direction estimates split by audit format.Forced-response and forced-choice tools over-correct, whereas generative and default-output tools remain stereotype-congruent.
- 1 Introduction: The study adds six deliberately weaker models as a positive control to test whether ranking failure reflects insufficient capability differences rather than tool disagreement.This control asks whether measurement recovers when the model panel spans real capability gaps.
2 Related work
Prior work shows that intrinsic fairness measures correlate poorly and that bias is conceptually fractured. This paper extends comparison to extrinsic, output-level audits under a controlled shared-model design.
- 2 Related work: Earlier studies find poor agreement among embedding-based fairness measures and emphasize that bias is a fractured concept across NLP.The cited literature provides precedent for measurement disagreement in intrinsic evaluation.
- 2 Related work: This study moves the comparison to extrinsic output-level tools targeted by regulation, fixing the model panel and construct to reduce access and panel confounds.It also adds reliability analysis to quantify and interpret tool disagreement.
- 2 Related work: Intrinsic benchmarks and holistic suites are excluded because a pooled-API design cannot accommodate their embedding access or heavy adapters.The comparison therefore focuses on standard extrinsic tools compatible with the common harness.
3 Method
The method reimplements ten extrinsic tools in a common harness over shared models and maps their outputs to comparable signed scores. It evaluates detection, direction, ranking, and reliability under explicit tool-inclusion rules.
- 3 Method: Ten tools probe occupational gender bias across ambiguous QA, comparative questions, generation, surveys, coreference, and forced-choice decisions.The primary analysis uses a common harness and pooled routing gateway for the frontier-model panel.
- 3 Method: Native outputs are harmonized into signed per-item scores, with positive indicating male-favoring or stereotype-congruent results and negative indicating female-favoring results.BBQ and BiasAsker retain magnitude-only scores because neutral answers prevent stable signed directions.
- 3 Method: Detection uses model-mean absolute net scores with bootstrap confidence intervals, direction uses model-mean net scores and sign-preserving item bootstraps, and ranking uses Kendall statistics.Reliability is estimated with split-half correlations and Spearman-Brown correction, alongside variance ratios and ICC.
- 3 Method: BBQ and BiasAsker are excluded from main detection, direction, and baseline ranking analyses, re-entering only ranking sensitivity analysis while remaining reported for reliability.The inclusion matrix makes these analysis-specific exclusions explicit.
- 3 Method: Reliability is measured across items, while the variance ratio and ICC quantify how much score variance reflects differences between models rather than within-model item variation.These measures operationalize whether each tool can reproduce model-level judgments.
4 Results
Across occupational gender bias, the tools reliably detect bias but fail to agree on model rankings. Adding weaker models restores some within-tool reliability without restoring cross-tool agreement, while bias direction varies systematically by audit format and the detection-ranking pattern largely replicates across constructs.
- 4.1 The tools detect bias: Every one of eight discriminating tools detects occupational gender bias, while BBQ and BiasAsker are saturated by mostly neutral frontier-model answers.Model-bootstrap 95% intervals exclude zero for the eight tools; saturation collapses between-model variance, with BBQ’s variance ratio at 0.005.
- 4.3 The tools disagree on how to rank models: Kendall’s W = 0.07 with p = 0.83, showing that eight discriminating tools do not rank frontier models beyond chance.The observed W is below the permutation-null 95th percentile of 0.23; mean pairwise τ = −0.05.
- 4.4 Why ranking fails: a positive control: Adding six weaker models raises reliability for several interpretable tools, but cross-tool ranking remains nonsignificant on both the full and frontier panels.Marked Personas rises from 0.39 to 0.71 and BiasLab from 0.11 to 0.22; ranking remains nonsignificant, including W = 0.13 on the sixteen-model panel and W = 0.16 in the matched six-tool check.
- 4.5 What the tools disagree about: direction: Bias direction splits by format: forced-choice decision tools are female-favoring, whereas spontaneous, default, and opinion-output tools remain male-stereotypical.BiasLab is −0.38 and Hiring −0.37, while Marked Personas is +0.57, WinoBias +0.27, and OpinionQA +0.19.
- 4.6 Does the pattern replicate? Age and socioeconomic status: The detection pattern mostly replicates on age and socioeconomic status, but ranking agreement remains absent for socioeconomic status and the apparent age agreement is fragile.Socioeconomic-status ranking is nonsignificant; age shows W = 0.32, p = 0.012, but this result uses seven models, has a wide interval, and fails the paper’s robustness concerns.
- 4.7 Two side findings: Claude Sonnet 4’s refusal of every Hiring item is treated as a separate behavioral category and excluded from that tool’s score and listwise ranking statistics.The refusal reflects declining to make hiring decisions from demographic proxies, leaving no comparable Hiring score for Claude.
5 Discussion
Across constructs, audits reliably detect bias, but cross-tool rankings fail because reliability depends on model separation while agreement is limited by different operationalized sub-constructs.
- 5 Discussion: The positive control separates two failures: within-tool reliability improves when models span capability gaps, whereas cross-tool ranking remains null because tools measure different channels.The authors favor construct multiplicity over shared noise, while acknowledging systematic uncorrelated item-sampling noise as an alternative.
- 5 Discussion: Across three constructs, audits agree that bias is present, but tools disagree on which model is worst and estimate direction differently by operationalization and response format.Socioeconomic-status results reproduce ranking failure, while age agreement disappears under the paper’s robustness rules.
- 5 Discussion: Direction aligns by format: forced-choice tools tend toward female-favoring corrections, while free-generation and coreference tools remain male-stereotypical.This structured directional agreement alongside rank disagreement supports distinct sub-constructs rather than one construct measured noisily.
- 5 Discussion: Two saturated tools cannot rank this converged frontier panel, but the paper treats saturation as panel-dependent rather than an intrinsic defect.The limitation concerns the current model panel, not a permanent verdict on those tools.
- 5 Discussion: Governance can use a well-chosen audit to detect bias and estimate direction within a stated operationalization, but one audit score does not support cross-model ranking.Default-output and forced-choice audits target different debiasing problems, so no single benchmark certifies both.
- 5 Discussion: For a target reliability of r*=0.8, projected item counts are about 248 for Marked Personas, 1,000 for BiasLab, and 38 for Hiring.These projections illustrate how current item banks constrain reliable ranking when models are highly similar.
6 Limitations
The main limitations concern reimplementation fidelity, proxy scorers, provenance, and incomplete positive-control coverage.
- 6 Limitations: All ten tools were rebuilt inside one common harness rather than run from their original codebases, so direct comparison with published scores remains unavailable.The authors identify validation against one or two published implementations as the first follow-up.
- 6 Limitations: Reference Letters and BOLD proxy original classifier-based metrics, while Reference Letters shows large item effects that cancel in aggregate and yields weak reliability.Removing Reference Letters does not change the ranking conclusion, but the authors rely on it lightly.
- 6 Limitations: BiasLab was in press rather than independently peer-reviewed, and Hiring was bespoke; Hiring also anchors the female-favoring direction result.The paper releases Hiring’s full specification for scrutiny or replacement.
- 6 Limitations: The reliability-recovery positive control excludes Hiring and WinoBias from the weak cohort, so recovery is shown only for the other tested tools.The observed recovery is consistent across tested tools but cannot be demonstrated directly for the two most reliable tools.
Reproducibility
The study releases its harness, model panels, raw responses, item-level scores, harmonization specification, and analysis code for recomputation.
- Reproducibility: The released materials include pinned panels, model identifiers and retrieval dates, harmonization specifications, raw responses, item-level scores, and analysis code.The repository’s consistency script recomputes every headline number directly from raw responses.
A Harmonization sensitivity
Ranking conclusions remain non-significant across five harmonization variants, with magnitude scoring providing the relevant test of which model is most biased.
- A Harmonization sensitivity: Raw signed scoring produces a positive mean τ of +0.08 because tools share a directional tilt, not because their rankings genuinely agree.The paper therefore distinguishes shared direction from concordance in model ordering.
- A Harmonization sensitivity: No harmonization variant yields significant ranking agreement, and magnitude scoring stays at or below chance for the model-ranking question.Table 4 tests each variant against its own permutation null, whose threshold changes with tool count.
- A Harmonization sensitivity: The table’s null percentiles use 20,000 Monte Carlo resamples, leaving approximately ±0.01 sampling noise in the second decimal.Variants A and B share the same eight-tool, nine-model null.