Source-linked AI summary
Whose Gold? Annotator-Pool Disagreement Is Large at the Item Level, and Hidden by Small Leaderboards
Anik Jha
TL;DR
Preference benchmarks leave the choice of annotator pool largely unexamined, despite its potential to affect measured preferences. This paper compares expert and crowd pools, quantifies how their perturbations affect rankings and LLM judges, and finds substantial item-level divergence alongside leaderboard invariance that does not generalize to larger arenas.
Problem
Preference benchmarks typically treat the choice between crowdworkers and domain experts as an implementation detail, leaving its effect on measured preferences underexamined.
Method
The paper compares expert and crowd labels on convention-free items, propagates pool differences through leaderboard perturbations, and evaluates three LLM judges against both pools.
Results
Expert and crowd pools reverse winners on 9.2% of MultiPref items and 8.5% of MT-Bench cells, while all three LLM judges agree more with crowds than experts.
Takeaways & Limitations
Six-model leaderboard invariance reflects its spacing rather than a general aggregation property, while larger leaderboards are more likely to change under pool substitution.
Takeaways & Limitations
The larger-leaderboard extrapolation assumes model-wise independent, approximately Gaussian pool perturbations and is checked directly only with six models.
Abstract
from arXiv · showhide
Preference benchmarks are built by hiring annotators, and the identity of those annotators is treated as an implementation detail. We measure what that detail buys. On the 2,885 MultiPref items where both pools are internally unanimous, so no tie-breaking convention is consulted at all, expert and crowd annotators assign a different majority label to 23.6% and name the opposite winner on 9.2%; on the 246 comparably unanimous MT-Bench cells, benchmark authors and recruited experts differ on 30.5% and reverse on 8.5%. Yet on both corpora the resulting model leaderboards are bit-identical: Kendall tau = 1.00 with zero of six models displaced. That invariance is far weaker evidence than it looks, and we quantify how weak. Switching pools moves a model's win rate by 1.9pp (SD), one adjacent pair in our own leaderboard sits 0.8pp apart and had a 38% chance of swapping, and an item-level bootstrap displaces at least one model in 28% of resamples. The observed zero is the common outcome, not a property of aggregation: on the same measured perturbation, a ten-model leaderboard is displaced with probability 0.86 and a twenty-model leaderboard with probability 0.9997. Reporting a six-model leaderboard is safe; the safety does not generalise, and everything that consumes labels per item is not safe at any size. We make the distinction precise, show that a widely used dataset's stated assumption of no intra-group annotator variability is false, and show that an LLM judge tracks the crowd pool over the expert pool on all three models we test, including one from a different vendor. All code, per-call outputs, and pre-registered decision rules will be released upon acceptance.
1 Introduction
The paper shows that annotator-pool choice produces substantial item-level disagreement, even though expert- and crowd-labeled leaderboards can remain unchanged. This distinction matters because per-item label consumers, including LLM judges, inherit the hiring decision.
- Item-level divergence: The stated assumption of no intra-group annotator variability does not survive direct comparison of ordinary crowdworkers and screened domain experts.MULTIPREF explicitly identifies that assumption and leaves its exploration for future work, which this paper undertakes.
- Item-level divergence: 23.6% and 30.5% of items change majority labels across pools in MULTIPREF and MT-BENCH, while winner reversals occur on 9.2% and 8.5%.These figures use items where no tie-breaking convention is needed; two MULTIPREF experts also disagree on 50.1% of items.
- Leaderboard versus item labels: τ = 1.00 on both corpora, with no model displaced, when the same models are ranked under expert-gold versus crowd-gold labels.The paper attributes this invariance to leaderboard spacing and argues that displacement should be expected for larger leaderboards.
- Leaderboard versus item labels: Per-item labels are consumed by reward-model training, active-annotation routing, and LLM-judge validation, so leaderboard safety does not establish safety for these uses.The paper measures the judge case and conjectures the others share the mechanism.
- LLM-judge validation: −6.9, −5.2 and −3.7pp: all three tested LLM judges agree more with the crowd majority than the expert majority, with confidence intervals excluding zero.Two judges come from one vendor and one from another, making judge–human agreement partly a consequence of the hiring decision.
2 Setup
The study contrasts crowd and expert judgments on two preference corpora using per-item majority labels and model mean credit, while avoiding tie-breaking by restricting its primary analysis to internally unanimous items. It also documents why comparable annotator-pool contrasts are unavailable in two additional corpora and pre-registers the leaderboard decision rule.
- Corpora: 10,461 MULTIPREF comparisons cover six models, with each item annotated twice by crowdworkers and twice by screened experts.MT-BENCH contributes 3,355 human votes across 292 cells rated by both benchmark authors and recruited experts.
- Corpora: PRISM and HELPSTEER2 cannot support annotator-pool contrasts because they lack repeated per-item annotations or annotator identities with pairwise preferences.Multi-annotated preference corpora retaining per-annotator labels are scarce.
- Aggregation: Each pool receives a majority label per item, and each model is scored by mean credit over the comparisons in which it appears.Rank comparisons use Kendall τ with paired bootstrap resamples drawing identical item indices for both label sources.
- Pre-registration: τ > 0.9 with a confidence-interval lower bound > 0.9 was the pre-registered rule for reporting the null and stopping without enlarging the model set.The reported outcome follows this versioned decision rule rather than a post-hoc rule.
- Tie handling: 2,885 items, or 27.6% of the corpus, remain after restricting to cases where both pools are internally unanimous, avoiding tie-breaking conventions.The two experts disagree 50.1% of the time, so pool-majority decisions otherwise depend on convention on roughly half the corpus.
3 Divergence at the item level, invariance in aggregate
Expert and crowd annotators frequently disagree at the item level, including across comparisons with different model-quality gaps, yet their aggregate six-model leaderboards remain identical. This cancellation does not show that item-level reversals are negligible for consumers that do not average labels.
- Item-level divergence: 48.1%, 50.3% and 51.4% majority divergence occurred across quality-gap tertiles, with Pearson r = +0.030, showing disagreement was flat across comparison difficulty.The tertiles span closely matched through widely separated model comparisons.
- Item-level divergence: One comparison in eleven had expert and crowd majorities naming opposite winners on the convention-free MULTIPREF subset.This is an item-level reversal rate that can cancel in an aggregate mean.
- Aggregate invariance: τ = 1.0000, with zero of six models displaced, because both leaderboards were identical on both corpora.The bootstrap confidence intervals were [0.867, 1.000] for MULTIPREF and [0.733, 1.000] for MT-BENCH.
- Implications: 9.2% item-level reversal remains consequential for consumers that do not take the mean, despite cancellation in aggregate.The passage distinguishes aggregate cancellation from the persistence of item-level reversals.
4 How much invariance is that, exactly?
A Kendall tau of 1.00 does not show that aggregation is robust: the measured annotator-pool perturbation could displace models, but the six-model leaderboard happened to survive. Displacement probability rises sharply with leaderboard size, reaching 0.9997 at 20 models.
- Where the invariance ends: τ = 1.00 licenses only the narrow conclusion that a six-model leaderboard with one large gap survived this perturbation, not that leaderboards are robust to annotator identity.The reassuring interpretation is available only at small K, and real arenas cluster models more tightly than the uniform-spacing estimates.
- The perturbation: 1.9pp SD: switching annotator pools shifts six models’ win rates by −2.1pp to +2.8pp, while adjacent-pair shift differences have 2.6pp SD.Any adjacent models within about 4.3pp have at least a 5% chance of swapping when the annotator pool changes.
- The perturbation: 28%: item-level resampling displaced at least one model on the six-model leaderboard, making the observed zero a modal but unreliable outcome.The 0.8pp adjacent pair had a 38% swapping probability and did not swap; the parametric estimate was 28.1%, versus 28.2% from the bootstrap.
- Where the invariance ends: 0.86: the probability of displacing at least one model rises at K = 10, compared with 0.28 at K = 6 and 0.9997 at K = 20.At K = 50, displacement is indistinguishable from one; these estimates hold the measured perturbation fixed and use uniform spacing across the observed win-rate span.
5 What inherits the divergence
LLM judges inherit annotator-pool divergence: all three tested models align more with crowd than expert majorities, including a model from a different vendor. The effect is robust across matched-power runs but decreases as judges become cleaner, while position sensitivity remains substantial.
- Interpretation: Judge agreement with humans partly reflects which annotator pool was hired when the judge tracks one pool more than the other.This makes reported human-agreement percentages partly statements about the hiring decision.
- LLM-judge alignment: All three models agree more with the crowd majority than the expert majority on MULTIPREF, including one from a different vendor and pretraining lineage.The models used both presentation orders and a schema-constrained five-point output matching the human scale.
- Replication and limitation: −3.2pp, CI [−7.3, +1.2] was observed in an initial half-size run of the second model, which failed the pre-registered replication bar before the effect appeared at matched power.The initial result is reported as an honest complication rather than treating the matched-power run as the first attempt.
- LLM-judge alignment: 6.9 → 5.2 → 3.7pp crowd lean accompanies order-swap reversal falling from 44.6% to 24.5% to 23.7% across the three judges.The cleaner the judge, the smaller the pool-alignment effect, contradicting the proposed label-noise attenuation explanation; the mechanism remains open.
- Measurement choices: 44.6%, 24.5% and 23.7% of items show direction reversal under order swaps, making position sensitivity a substantial measurement choice in the judge arm.These effects differ by a factor of 1.8 between two judges and exceed the pool effect discussed in the section.
6 What to do instead
Reliable preference-benchmark interpretation requires reporting the annotator pool, separating leaderboard stability from per-item reliability, counterbalancing judge order, and preserving per-annotator labels. These practices are necessary because pool identity affects reference labels, order can reverse judgments 24.5–44.6% of the time, and discarded annotator identities prevent analysis in two of four corpora.
- Reporting: Report the annotator pool: an LLM-judge agreement figure without its reference-label pool is under-specified by roughly the reported effect size.The reference labels’ defining annotator pool is necessary to interpret agreement figures.
- Interpretation: Do not infer per-item reliability from leaderboard stability, because the data separates these quantities by a wide margin.At six models, leaderboard stability largely reflects model spacing; §4 provides arithmetic for checking whether a board is in the safe regime.
- Measurement design: 24.5–44.6% direction reversal makes a single-order judge measurement on this task close to a coin flip.Judge order should therefore be counterbalanced in every measurement.
- Data retention: Two of four examined corpora cannot support the analysis because they discarded annotator identity.Per-annotator labels must be retained to enable this analysis.
7 Limitations
The study’s aggregate invariance is demonstrated only on six-model leaderboards, so its extrapolation to larger arenas remains conditional. That extrapolation assumes independent, approximately Gaussian pool perturbations and requires validation on corpora with more models.
- Scope of evaluation: Six models: both corpora rank only six models, limiting what the observed aggregate invariance can establish.The spacing of this leaderboard, rather than aggregation generally, determines the result.
- Extrapolation assumptions: Independent, approximately Gaussian perturbations: Table 3’s extrapolation relies on these assumptions across models.The model reproduces the K = 6 bootstrap closely, but the available validation set contains only six models.
- Future validation: More-model corpora: testing larger arenas directly is needed to validate whether the extrapolation transfers.The passage identifies a corpus with more models as the obvious way to test this properly.
8 Related work
Prior work treats annotator disagreement as informative and models annotator confusion in aggregation. This paper instead asks whether annotator population changes evaluation outcomes and measures LLM judges against each population separately.
- Annotator disagreement: Disagreement-aware evaluation treats annotator disagreement as signal rather than noise.This perspective is attributed to Xu and Jurgens (2026).
- Annotator disagreement: Other methods propose aggregation procedures that model annotator confusion when ranking systems.Bonagiri et al. (2026) is cited for this approach.
- Annotator populations and LLM judges: This paper studies whether the choice of annotator population changes evaluation answers and separately compares LLM judges with each population.The comparison is framed as a prior question to disagreement-aware aggregation.