Source-linked AI summary
What Does an LLM-Agent Leaderboard Rank Actually Compare?
Wei-Jung Huang
TL;DR
The paper asks when a leaderboard score justifies treating one LLM agent as superior to another despite differences in targets, labels, release detail, and cost rules. It introduces an estimand-aware pairwise procedure that checks support and evaluates uncertainty and practical margins. Across public benchmarks, close comparisons often remain unresolved, while proxy labels and utility rules can change the selected system.
Problem
Public leaderboard scores may not support broader pairwise superiority claims when task mixtures, label sources, release details, or cost rules differ.
Method
The procedure states the comparison target and measurement source, checks common support, and evaluates the supported gap using an uncertainty rule and practical margin.
Results
Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are often unresolved, while proxy labels, repeated-attempt rules, and utility rules can change the score or selected system.
Takeaways & Limitations
A leaderboard score summarizes one released evaluation; a fine-grained superiority claim additionally requires its estimand and decision rule.
Takeaways & Limitations
Public releases support pairwise conclusions only for recorded variables and outcomes and generally do not identify causal effects of repositories, benchmarks, scaffolds, modes, or verifiers.
Abstract
from arXiv · showhide
An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent. Public evaluation logs may not support that conclusion when systems differ in task mixture, label source, release detail, or cost rule. We study what leaderboard scores estimate and when they justify pairwise superiority conclusions. Our estimand-aware pairwise procedure states the comparison target and measurement source, checks common support, and evaluates the supported difference using a stated uncertainty rule and practical margin. Controlled checks evaluate the decision labels under known finite-sample conditions and show why uncertainty must be included when judging sensitivity to target reweighting. Across SWE-bench, AgentRewardBench, and tau2-bench, close rank differences are often unresolved; proxy labels and utility rules can also change which system is selected. DataAgentBench and Open Agent show what remains estimable from coarser public records. A leaderboard score summarizes a released evaluation, whereas a fine-grained superiority claim additionally depends on the estimand and uncertainty rule used to interpret the difference.
I. INTRODUCTION
Leaderboard ranks summarize particular evaluation designs, not automatically general agent superiority. The paper proposes an estimand-aware procedure that scopes pairwise claims by target, labels, support, uncertainty, and decision rule.
- Motivation: Leaderboard rows summarize configurations that differ in scaffolds, tools, prompts, stopping rules, label sources, and resource budgets.A displayed ordering therefore reflects a particular evaluation design rather than a general ordering of agents.
- Motivation: Expert labels selected Claude 3.7 Sonnet on AgentRewardBench, whereas GPT-4o-mini judge labels selected GPT-4o.The example shows that changing the label source can change the selected system.
- Empirical motivation: On SWE-bench Verified, 5 of 45 top-10 pairs were stable, 1 was target-sensitive, and 39 were underpowered; simultaneous intervals resolved none.The comparisons used descriptive pointwise 95 percent intervals and a 2-point decision rule, while a max-deviation simultaneous interval resolved none of the 45.
- Empirical motivation: On tau2-bench, pass@4 raised scores by 17.7 to 20.4 points relative to pass@1, while cost changed which system a utility rule selected.Repeated-attempt protocols and utility rules can alter both the measured score and the selected system.
- Conclusion: A released score answers one evaluation question, whereas a broader pairwise ordering requires the score estimand and comparison conditions to match.The paper frames the distinction as a difference between a released evaluation and a scoped superiority claim.
- Procedure: The procedure records the evaluated unit and target, checks common support, and evaluates the supported gap with a stated uncertainty rule and practical margin.It reports a supported winner only when the evidence permits one; otherwise it identifies the limiting condition.
- Evaluation: Controlled analyses test label behavior under known scenarios before applying the same procedure to SWE-bench, AgentRewardBench, and tau2-bench.The study also uses public records from DataAgentBench and Open Agent to examine what remains estimable from coarser releases.
II. RELATED WORK
Related work supplies tools for explicit targets, uncertainty-aware comparison, benchmark documentation, and utility tradeoffs. This paper uses those foundations to identify which pairwise conclusions remain estimable from public agent logs.
- Benchmark documentation: Agent benchmarks and evaluation documentation define tasks, environments, harnesses, verifiers, and public reporting formats for interactive systems.Prior work also emphasizes cost control, reproducibility, benchmark design, and structured result documentation.
- Targets and labels: Survey sampling, post-stratification, and causal-inference work make target populations and estimands explicit before estimation or identification.Calibration weighting, entropy balancing, hierarchical models, and item-response models provide stronger estimators when their assumptions and data support hold.
- Robustness and inference: Benchmark perturbation, prompt-distribution, and uncertainty analyses show that conclusions can change under altered evaluation conditions.Classifier-comparison and multiple-comparison methods treat rank differences as inferential claims rather than point estimates alone.
- Utility and public-data limits: Multi-criteria decision analysis formalizes utility tradeoffs, while limited public logs constrain full causal or off-policy analyses.The paper therefore focuses on score targets and pairwise conclusions that remain estimable from released records.
III. A PAIRWISE COMPARISON PROCEDURE
The procedure makes leaderboard comparisons estimand-aware by defining the evaluated configuration, target, measurement source, and support conditions before judging pairwise superiority. It separates observed-mix, target-standardized, verifier-based, and utility-based scores, which answer different questions.
- Comparison target: The evaluated unit is a full agent configuration, including the base model, scaffold, tools, action space, prompt policy, stopping rule, verifier, and resource budget.A model-only claim requires additional design support across multiple scaffolds and models.
- Comparison target: An estimand specifies the quantity a score targets, such as success on the observed mix, a stated target population, a label source, or a resource rule.The same observed run can support different score definitions under different target populations or utility rules.
- Comparison target: An ordering under one estimand should not be read as an ordering under another because observed success, target-standardized success, weak-verifier success, and utility answer different questions.Capability, measurement, and utility claims therefore require distinct target and evidence specifications.
- Support: The procedure checks whether every compared configuration has observations in the target’s supported cells before standardizing scores or differences.If the target assigns mass outside common support, the comparison is unsupported without modeling assumptions or must use a restricted target.
- Standardization: Direct standardization reweights cell-level empirical success rates by target weights to report how scores and pairwise differences change under the stated target.This reweighting describes target-standardized comparisons rather than causal effects of cells.
C. When Is a Score Difference Large Enough?
The procedure treats a score difference as a supported practical superiority claim only when its direction is supported by uncertainty and its magnitude meets the stated margin. It also labels comparisons blocked by unsupported targets, verifier disagreement, insufficient precision, or target sensitivity.
- Decision rule: A practical winner requires an interval excluding zero in its direction and a point estimate at least as large as the practical margin.For a, the rule is L_ab > 0 and Δ̂_ab ≥ m; for b, U_ab < 0 and Δ̂_ab ≤ −m.
- Decision labels: A comparison is underpowered when its interval includes zero or its target-standardized gap is smaller than the stated practical margin.This label does not imply that more samples would change the conclusion.
- Target sensitivity: A comparison is target-sensitive when a supported target reverses the displayed ordering or uncertainty-supported movement from the released mix exceeds the practical margin.The analysis asks both about one specified alternative target and about the minimum reweighting needed to change the result.
- Practical margin: The default reporting margin is 2 percentage points when no release-specific cost or switching rule is available.A release-specific rule takes precedence when one exists.
- Uncertainty: Bootstrap intervals resample tasks or reported cells according to repeated structure, while aggregate-only datasets receive grouped stability checks rather than item-level uncertainty claims.Named-pair claims may use pointwise intervals; tier-wide or winner claims require a declared comparison family and error-control target.
- Decision labels: The ordered labels are unsupported, verifier-sensitive, underpowered, target-sensitive, and stable, with the earliest blocking issue reported compactly.All diagnostic flags remain available even when only one ordered label is displayed.
IV. CONTROLLED BEHAVIOR CHECKS
Controlled reference cases are used before public leaderboard analysis because the true ordering is unknown. The checks assess label recovery, target standardization, weak-verifier calibration, and a task-conditioned null reference.
- Controlled checks: Controlled simulations assess whether the procedure recovers known decision states and whether standardization and weak-verifier calibration reduce error when their assumptions hold.An AgentRewardBench permutation analysis provides a task-conditioned null reference for observed balanced gaps.
A. Decision-Label Recovery
The decision labels are evaluated under finite-sample scenarios with known success probabilities, target mixes, support patterns, verifier behavior, and margins. Uncertainty-aware target-sensitivity labeling substantially improves recovery of a stable scenario over point-estimate movement alone.
- Simulation design: The label-recovery simulation uses two configurations, four evaluation cells, 300 replicates per scenario, and 300 bootstrap draws per replicate.Each scenario fixes the true cell probabilities, observed and target mixes, support pattern, verifier behavior, and practical margin before outcomes are sampled.
- Label recovery: The modal decision label matches the known state in every simulated scenario, although finite sampling still produces occasional deviations.Table I reports Wilson 95 percent intervals over replicates.
- Target sensitivity: 84.3 percent of replicates recovered the stable scenario under point-estimate movement alone, with target sensitivity overcalled in 47 of 300 runs.This simpler rule is compared against the uncertainty-aware rule.
- Target sensitivity: 99.7 percent of replicates recovered the stable scenario under the uncertainty-aware rule.The empirical analysis therefore uses uncertainty-supported movement when item-level or cell-level uncertainty is available.
B. Known-Target Estimation Check
The known-target simulation evaluates whether standardization reduces estimation error, while permutation and perturbation checks assess decision behavior under finite-sample uncertainty and target variation.
- Direct standardization reduced mean absolute error from 0.0184 for the raw observed expert score to 0.0128.The simulation used four configurations across assistant, web, visual, and work conditions, with equal benchmark target weights.
- The AgentRewardBench permutation test found an observed maximum absolute balanced pairwise gap of 0.1489 versus a permutation 95th percentile of 0.0727.The plus-one corrected empirical p value for the maximum gap was below 0.001.
- The count of practical gaps at least 2 percentage points exceeded the permutation reference, with plus-one corrected empirical p value 0.010.
- Claude remained the top agent in all 1,000 target-weight perturbation draws, but the 10th percentile of the top gap was 0.0165.The top position was stable near the equal-benchmark target while some pairwise margins remained small.
V. EMPIRICAL EVIDENCE
Across detailed public releases, target reweighting can shift scores and point estimates, but uncertainty often prevents fine-grained superiority claims from being resolved.
- Repository-balanced scoring changed practical conclusions on SWE-bench: a raw tie at 0.792 separated into balanced scores of 0.768 and 0.780.The corresponding bootstrap intervals were [0.690, 0.833] and [0.694, 0.846], and the maximum absolute shift across submissions was 12.0 points.
- Among 45 SWE-bench pairs, 5 were stable, 1 was target-sensitive, and 39 were underpowered under a 2-point practical-margin rule.The labels used descriptive pointwise bootstrap intervals and were conditional on the observed raw top 10.
- Equal-repository reweighting moved 38 of 45 gaps by at least 1 point, 31 by at least 2 points, and 7 by at least 5 points.A max-deviation simultaneous bootstrap interval excluded zero for none of the 45 pairs.
- The public evidence rarely resolves close unadjusted SWE-bench orderings within the observed raw top 10.
C. AgentRewardBench: Coverage Limits the Target and Label Source Changes the Winner
AgentRewardBench shows that common support constrains the target and that proxy label sources can alter both score levels and the selected winner.
- The full all-agent target is unsupported because Llama has no VisualWebArena observations, so the common-support target excludes VisualWebArena.Under common support, ranks did not change, but every score decreased.
- Under the common-support target, Claude 3.7 Sonnet fell from 0.339 raw success to 0.300 benchmark-balanced success, while Llama fell from 0.181 to 0.154.The largest absolute shift was 5.1 points.
- Four of six common-support pairwise comparisons were stable and two were underpowered under the uncertainty-aware rule.Benchmark mixture changed score levels without reversing the supported order.
- Expert labels selected Claude 3.7 Sonnet, whereas GPT-4o-mini labels selected GPT-4o; each weak verifier reversed one of six pairwise signs.The functional evaluator underestimated expert success by 8.6 points, while GPT-4o-mini overestimated it by 7.9 points.
- Calibration used paired weak-verifier and expert labels, and releases without such pairs permit only comparisons between label sources.
D. tau2-bench: Protocol and Utility Choices Change Selections
tau2-bench demonstrates that repeated-attempt protocols, task representation, cost rules, and release granularity each constrain or change agent comparisons.
- Equal-domain reweighting shifted tau2-bench scores by at most 2.2 points, yet all six pairwise comparisons remained underpowered.The main analysis used four configurations with common support over airline, retail, and telecom.
- Averaging pass@4 over domains raised common-support main-mode scores by 17.7 to 20.4 points relative to pass@1.The largest pass@4 lift across all domain settings was 24.6 points, while o4-mini ranged from 0.421 to 0.989 pass@1 across four telecom modes.
- The top system remained unchanged after logistic adjustment, with equal-domain scores moving by at most 1.58 points relative to direct standardization.Sensitivity to task distribution depended more strongly on how distance between mixed-type tasks was defined.
- A cost-aware utility rule switched the point-estimate winner from Claude 3.7 Sonnet to o4-mini at λ = 0.0514 success-rate units per USD.The utility was success minus λ times agent-side cost, with duration weight µ = 0.
- Coarser releases still support some aggregate comparisons, but they do not support item-level uncertainty or verifier calibration in Open Agent.DataAgentBench retained PromptQL Claude Opus 4.6 as top under equal-dataset scoring despite score shifts up to 9.2 points and 10 of 28 pairs remaining underpowered.
VI. DISCUSSION AND LIMITATIONS
Public releases support pairwise conclusions only for recorded variables and outcomes, while sensitivity and multiplicity choices determine the scope of any reported winner. Accordingly, leaderboard authors should expose configuration, target, labels, support, uncertainty, and practical-significance rules rather than silently collapsing decision dimensions into one ordering.
- Reporting implications: Leaderboard authors should report the evaluated configuration, score-defining population and labels, support, uncertainty, and practical-significance rules.These disclosures make the conditions for pairwise superiority visible.
- Reporting implications: Cost, latency, repeats, and scaffolds should remain separate decision dimensions rather than being silently folded into one ordering.
- Scope of supported conclusions: Public releases support pairwise conclusions only for the variables and outcomes they record, not causal effects of repositories, benchmarks, scaffolds, modes, or verifiers.AgentRewardBench permits calibration through paired expert and weak labels; most releases permit only comparisons between label sources.
- Scope of supported conclusions: Sensitivity and multiplicity choices determine the scope of any reported winner.Pointwise intervals suit named-pair claims, whereas family-level winners require a declared comparison family and error-control rule.
- Overall boundary: Close ranks often do not support a pairwise winner once the target, common support, and uncertainty are specified.A leaderboard score summarizes one released evaluation; pairwise superiority additionally requires its estimand and decision rule.