Source-linked AI summary
The Leaderboard Illusion
Shivalika Singh, Yiyang Nan, Alex Wang, Daniel D'Souza, Sayash Kapoor, Ahmet Üstün, Sanmi Koyejo, Yuntian Deng, Shayne Longpre, Noah A. Smith, Beyza Ermis, Marzieh Fadaee, Sara Hooker
TL;DR
The paper asks whether Chatbot Arena’s growing role as a model benchmark is undermined by private testing, selective disclosure, and unequal access to evaluation data. It audits 2M battles across 243 models and 42 providers, combining analyses, simulations, and controlled fine-tuning experiments. The results indicate biased rankings, substantial provider data asymmetries, and Arena-specific overfitting, motivating reforms for more transparent and fair benchmarking.
Problem
Chatbot Arena has become a dominant model-comparison benchmark, but its evaluation policies and provider practices may allow leaderboard performance to be optimized rather than reflect general model quality.
Method
The paper audits 2M battles across 243 models and 42 providers, analyzes private testing, sampling, and deprecation, and uses simulations plus controlled fine-tuning experiments.
Results
The study finds that selective disclosure and unequal sampling and deprecation distort rankings, while increasing Arena data from 0% to 70% raises ArenaHard win rate from 23.5% to 49.9%.
Takeaways & Limitations
Chatbot Arena requires transparent score disclosure, limits on private variants, equitable removals and sampling, and public deprecation reporting to restore reliable benchmarking.
Takeaways & Limitations
The study lacks comprehensive raw Arena data and its scraped sample covers only January–March 2025, potentially underestimating private variants for providers with fewer launches.
Abstract
from arXiv · showhide
Measuring progress is fundamental to the advancement of any scientific field. As benchmarks play an increasingly central role, they also grow more susceptible to distortion. Chatbot Arena has emerged as the go-to leaderboard for ranking the most capable AI systems. Yet, in this work we identify systematic issues that have resulted in a distorted playing field. We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. We establish that the ability of these providers to choose the best score leads to biased Arena scores due to selective disclosure of performance results. At an extreme, we identify 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release. We also establish that proprietary closed models are sampled at higher rates (number of battles) and have fewer models removed from the arena than open-weight and open-source alternatives. Both these policies lead to large data access asymmetries over time. Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data. We show that access to Chatbot Arena data yields substantial benefits; even limited additional data can result in relative performance gains of up to 112% on the arena distribution, based on our conservative estimates. Together, these dynamics result in overfitting to Arena-specific dynamics rather than general model quality. The Arena builds on the substantial efforts of both the organizers and an open community that maintains this valuable evaluation platform. We offer actionable recommendations to reform the Chatbot Arena's evaluation framework and promote fairer, more transparent benchmarking for the field
1 Introduction
Chatbot Arena has become a central benchmark, but undisclosed private testing, selective score disclosure, and unequal data access can distort its rankings. The paper audits these mechanisms and proposes reforms to improve fairness, transparency, and reliability.
- 1 Introduction: Chatbot Arena’s growing influence creates risks that providers optimize for leaderboard performance rather than general model quality.The paper frames this risk through Goodhart’s Law and argues that preferential policies and provider behavior amplify gamification.
- 1 Introduction: The audit combines 2M battles covering 243 models and 42 providers to examine private testing, sampling, and model deprecation.The analysis groups providers into proprietary, open-weight, and open-source categories to compare access and treatment.
- 1 Introduction: Providers can test many private variants and publicly release only the best-performing checkpoint, which systematically skews Arena ratings upward.The paper reports as many as 27 privately tested Meta models before the Llama 4 release and demonstrates the effect with experiments and simulations.
- 1 Introduction: Proprietary providers receive substantially more Arena data because of unequal sampling rates and deprecation policies, creating persistent access asymmetries.Google and OpenAI received estimated shares of 19.2% and 20.4% of Arena data, while 41 fully open-source models collectively received 8.9%.
- 1 Introduction: The paper recommends permanent score disclosure, transparent private-testing limits, equalized removals, fair sampling, and public reporting of deprecated models.The recommendations target selective disclosure, excessive private testing, unequal model removals, provider-biased sampling, and silent deprecations.
2 Overview of Methodology
The paper audits Chatbot Arena using multiple datasets and examines how private testing, data access, overfitting, and deprecation affect leaderboard reliability. It also explains the Arena Score and the Bradley–Terry assumptions underlying it.
- Research questions: The study asks how private testing, provider data access, data-driven overfitting, and model deprecation affect Arena rankings.These questions structure the paper’s analyses of selection bias, data asymmetries, score impacts, and reliability.
- Methodological scope: 2M battles across 243 models and 42 providers form the paper’s empirical basis for auditing Chatbot Arena trends.The analysis groups models into proprietary, open-weight, and open-source categories using reported licenses.
- Chatbot Arena: Chatbot Arena ranks language models through human pairwise comparisons of anonymous responses.Users select which of two model responses performs better or declare a tie.
- Arena Score: The Arena Score uses a normalized Bradley–Terry model rather than averaging win rates.Bradley–Terry estimates relative skill from pairwise comparisons while accounting for opponent strength.
- Model assumptions: Bradley–Terry assumes pairwise comparisons come from an unbiased sampling process.The paper uses this assumption as context for analyzing how Arena policies can affect score reliability.
3 Results: Impact of Private Testing and Selective Retraction on Arena Scores
The results show that private multi-variant testing and selective disclosure can systematically advantage participating providers. Simulations and Arena experiments indicate that choosing the best observed submission distorts scores and rankings, including when model variants are identical or only marginally different.
- Simulated experiments: Testing 10 variants increased the maximum identified Arena Score by approximately 100 points in simulation.The simulated lift arises because more submissions increase the chance of finding a high-scoring observed variant.
- Real-world experiments: Real-Arena experiments found large gains from submitting multiple variants, including two identical Aya-Vision-8B checkpoints.Identical-checkpoint experiments provide a conservative lower bound because score differences cannot be attributed to model-quality differences.
- Observed private testing: 27 private Meta models were observed in March 2025, while Meta, Google, and Amazon were key beneficiaries of undisclosed private testing.Meta and Google had 27 and 10 tracked private models, respectively, during January–March 2025; the estimate excludes specialized leaderboards.
- Interpretation: Selecting and publicly releasing only the top-scoring variant violates Bradley–Terry’s unbiased-sampling assumption and inflates leaderboard rankings.The best-of-N strategy converts multiple private estimates into an extreme-value submission rather than a single unbiased skill estimate.
- Selection bias: E[β̂Best] > E[β̂k] for every submitted variant when non-degenerate score variation and N ≥ 2 hold.Selecting the best observed score creates upward bias because each estimate is affected by finite match-sampling fluctuations.
- Leaderboard effects: A weaker model family can surpass a stronger family when the former uses best-of-N private testing and the latter makes one public submission.The simulation shows that Family B can rank lower despite having a generally stronger model pool.
4 Results: Impact of Data Access Asymmetries on Arena Scores
Unequal private testing, sampling, deprecation, and API access give some providers substantially more Arena data than others. Experiments show that additional Arena data can sharply improve Arena-specific performance while providing limited broader-task benefits.
- 62.8% of Arena data goes to OpenAI, Google, Meta, and Anthropic collectively, 68 times the combined share of named academic and nonprofit labs.
- 112% estimated relative performance gains on ArenaHard arise from incorporating Arena data, while benefits on other tasks remain limited.
- Private variants, provider sampling rates, model deprecation, and API hosting determine how much Arena data each provider receives.
- Arena prompts shift over time and include exact or near duplicates, creating opportunities for models trained on Arena data to fit its evolving distribution.
- 23.5% to 49.9%: increasing arena-mix from 0% to 70% more than doubles win-rate against Llama-3.1-8B-Instruct on ArenaHard.
- MMLU performance changes from 66.5% at 0_arena to 65.9% at 70_arena, indicating that Arena gains are highly evaluation-specific.
5 Results: Impact of Model Deprecation on Arena Scores
Model deprecation undermines Arena ranking reliability when task distributions change or comparison graphs become sparse. The study’s simulations and observations show that uneven retirement can distort rankings and disadvantage some model groups.
- Deprecation under a changing task distribution produces unreliable rankings by violating assumptions behind the Bradley-Terry model.
- A sparse or disconnected comparison graph yields inaccurate skill estimates, whereas a dense graph aligns rankings more closely with true skill.
- 87.8% of open-weight and 89% of open-source models were deprecated, compared with 80% of proprietary models.
- 205 models were silently deprecated by near-zero sampling, while 47 models were publicly listed as deprecated.
- The Bradley-Terry model requires stable evaluation conditions and a sufficiently connected comparison network for transitive ranking inference.
6 Recommendations and Guidelines for Improving Leaderboards
The paper recommends transparent, auditable controls for private testing, score retraction, sampling, and model deprecation. These measures aim to reduce provider-type asymmetries and improve confidence in leaderboard dynamics.
- Score retraction should be prohibited because submitting only the best privately tested variant enables selective disclosure and can inflate leaderboard results.
- Private variants should be capped and disclosed at the provider level to limit unfair advantages from repeated pre-release testing.
- Deprecation criteria are currently difficult to audit because terms, thresholds, conjunctions, and pricing conditions lack precise definitions.
- Retiring the bottom 30th percentile within each proprietary, open-weight, and open-source category would preserve balance across provider types.
- Sampling should be made fairer across provider types, with periodic reporting on active sampling procedures and rates.
- Public reporting should cover all tested models, deprecations, sampling rates, and model pairings to enable community oversight.
7 Limitations
The study identifies limitations in its data access, observation period, training-data coverage, and provider attribution. These constraints bound how completely it can characterize private testing and Arena manipulation.
- Data access: Arena preprocessing and withheld private battles prevent access to the original, comprehensive raw dataset.This limits investigation of adversarial voting and related reliability concerns.
- Temporal coverage: The scraped-random-sample covers only January–March 2025, coinciding with Meta’s Llama 4 launch.Provider counts with fewer launches during this period may be underestimated.
- Training-data coverage: The training experiments use only a fraction of the data believed available to some proprietary providers.The authors estimate proprietary models may have been trained on 5 to 10 times more data than used in the experiments.
- Attribution: Private-model attribution relies on self-identification because anonymous model identities are not publicly disclosed.The authors describe this proxy as reasonable but inherently approximate.
8 Related Work
Related work positions benchmarks as influential but imperfect instruments shaped by their design, data, and incentives. It also situates Chatbot Arena within human-voting evaluation and its known reliability challenges.
- Benchmarking and progress: Benchmarks shape research priorities and incentives, but their assumptions and dependencies can affect reported progress.The related literature argues that benchmarks are rarely impartial.
- Benchmark failure modes: Static leaderboards are vulnerable to data contamination and implicit overfitting.Prior work identifies these risks across broad task-based leaderboards.
- Benchmark standardization: Inconsistent metrics and task definitions complicate meaningful comparisons across benchmarks.Related work also notes that leaderboard accuracy can neglect compactness and fairness.
- Data quality and reproducibility: Evaluation reliability is threatened by label errors, complex data streams, and limited reproducibility.These problems can affect result consistency even on seemingly simple tasks.
- Real-world validity: Benchmarks may diverge from practical utility because strong test performance does not guarantee real-world performance.This disconnect is especially concerning as benchmarks rapidly saturate.
- Human voting: Human-voting benchmarks capture nuanced qualities such as coherence, harmlessness, and readability through real-world prompts and feedback.Such preference data also supports alignment methods including RLHF.
- Human-voting limitations: Live voting benchmarks employ safeguards but still face evaluation challenges that this paper does not address.The related work distinguishes existing protections from unresolved reliability concerns.
9 Conclusion
The conclusion argues that Arena practices and provider behavior have compromised ranking reliability, while emphasizing the platform’s substantial community value. It recommends transparent, uniform reforms to restore fairness and trust.
- Interpretation: The authors acknowledge that Arena’s popularity and visibility may have contributed to the gradual emergence of its systematic issues.They also recognize the organizers’ substantial effort and the platform’s community benefits.
- Recommendations: The authors recommend banning provider-controlled score disclosure, limiting private variants uniformly, and using transparent removal and sampling policies.They argue these changes would reduce ranking distortions and prevent benefits from concentrating among a few providers.
- Platform context: Chatbot Arena grew from a multi-university collaboration into a standalone live evaluation platform under LMArena.The platform attracted millions of participants and collected over 3 million votes.
- Platform context: Users submit prompts, compare anonymous model responses, and vote, producing ratings through Online Elo and Bradley-Terry methods.The ranking process converts human preferences into model ratings.
- Ranking model: Bradley-Terry estimates latent model strengths from pairwise outcomes using maximum likelihood and cross-entropy loss.The estimated coefficients can be transformed into Elo-like ratings.
- Ranking model: Arena rankings incorporate confidence intervals, so overlapping intervals make model ordering less certain.This adds statistical nuance beyond a model’s Arena Score alone.
- Selection bias: Selecting the best of multiple noisy estimates produces a strictly positive selection bias when estimates are non-degenerate.The result formalizes why reporting only the top submission overestimates typical performance.
- Selection bias: The selection-bias result becomes equality only when the estimate variance is zero.In that case, the distribution is constant and all submitted estimates are identical.
D Data sources
The study combines historical battles, prompts, leaderboard snapshots, provider-shared data, and scraped battles to analyze Arena distributions, access, deprecations, and private testing. Scraping also identifies anonymous private models through de-anonymizing prompts and model responses.
- Dataset overview: The analysis combines multiple data sources covering 2M battles and 243 models across 42 providers.These sources support analyses of Arena trends and provider access.
- Historical battles: The historical-battles dataset contains 1.8 million battles from April 2023 to January 2025.It combines publicly released and provider-shared battle data.
- API prompts: API prompts contribute 567,319 entries because most historical battles lack prompts and released datasets are deduplicated.These prompts support analysis of duplicates and near-duplicates.
- Leaderboard statistics: Leaderboard-statistics snapshots track ratings, rankings, and battle counts over time from January 9, 2024 to April 23, 2025.They support comparisons of provider data access and model deprecations.
- Scraping procedure: De-anonymizing prompts were designed to reveal model identities while follow-up questions minimized interference with leaderboard rankings.Votes exposing identities are discarded by Chatbot Arena, and the researchers used ties as a precaution.
- Battle subsets: The historical data includes public battles and 43,729 proprietary battles involving Command and Aya models.The proprietary subset was shared under a policy allowing providers to request 20% of data involving their models.
- Scraped samples: The scraping process collected 5.8K battles from January–March 2025 and about 500 Vision-leaderboard samples.The Vision samples helped identify 35 private vision models.
- Private-model identification: The researchers identified 64 private models from 10 providers and captured 14 additional private models whose providers could not be inferred.Provider attribution used the models’ responses to identity prompts.
E.4 Assignment of Private Variants to Providers
The paper assigns unidentified private variants to providers by matching their disclosed self-identifications and response patterns. The collected examples include variants attributed to Meta, Google, OpenAI, and Cohere.
- Provider assignment: Table 4 lists private models captured in scraped-random-sample or scraped-vision-sample data, with identity-revealing response counts and examples.The table cautions that some private models may have appeared in more battles than the captured responses indicate.
- Meta: Meta-associated variants repeatedly identify themselves as Llama or as trained by Meta.Examples include kronus, polus, momentum, and anonymous-engine-2, alongside many additional Llama variants.
- Meta: Several Meta-associated responses explicitly expand LLaMA as “Large Language Model Meta AI.”This wording appears across anonymous-engine-2, unicorn-engine variants, and frost.
- Meta: Other Meta-associated variants use shorter labels such as “Llama, Meta,” “AI Assistant, Meta,” or “Meta trained me.”These include aurora, helix, prosperity, raze, solaris, spectra, toi, vega, and zax.
- Google: Google-associated variants identify themselves as large language models trained by Google.The examples include gemini-test, enigma, goblin, phantom, gremlin, specter, centaur, moonhowler, and nebula.
- Google: One Google-associated response identifies Gemma as an open-weights model trained by Google DeepMind.The response also describes Gemma as publicly accessible and able to accept text and images as inputs.
- OpenAI: OpenAI-associated variants identify themselves as ChatGPT trained by OpenAI.The examples include anonymous-chatbot and additional responses without visible model names.
- Cohere: Cohere-associated variants identify themselves as Command, a large language model developed or trained by Cohere.The examples include grapefruit-polar-bear, sandwich-ping-pong, and cohort-chowder.
G Data Access Estimation for Different Providers
The paper estimates that Chatbot Arena data access is highly uneven because proprietary providers collectively receive far more model API calls than the broader research community.
- Data access estimation: 3M user votes generated approximately 6M model API calls because each battle contains two models.Figure 4 represents roughly 5K API calls per square.
- Data access estimation: Proprietary providers collectively access considerably more Arena data than the broader research community.The research community receives only a fraction of the model API calls represented in the estimate.
H Analysis of Prompt Repetitions in Arena Data
Arena prompts show substantial repetition within and across months, while private best-of-N testing can inflate leaderboard scores through selection of the strongest observed variant. The inflation is larger for heterogeneous variants and does not disappear when sampling noise vanishes.
- H Analysis of Prompt Repetitions in Arena Data: Within-month duplication rates are generally high under both exact string matching and text-embedding similarity.The result indicates numerous repeated prompts within individual months.
- H Analysis of Prompt Repetitions in Arena Data: Substantial cross-month duplication reveals recurring prompt patterns or frequently asked questions that simple analysis can identify.Figure 16 reports both near-duplicate or highly similar prompts and exact matches across months.
- I Simulation for Expected Lift from Private Testing: 50 Arena Score points is the simulated lift from testing 20 non-identical private variants.The paper treats heterogeneous variants as the more realistic prerelease scenario.
- I.1 Background: A provider can train N private variants, test them on a hidden Arena fork, and publicly submit only the highest-scoring variant.Selecting the maximum of noisy measurements creates extreme-value bias.
- Asymptotics: As n increases, identical-variant selection bias tends to zero because sampling noise decreases proportionally to 1/√n.The paper therefore states that identical-checkpoint bias eventually disappears.
- I.3.1 Extreme-value uplift: Heterogeneous-variant selection bias remains positive as n approaches infinity because true skill variation remains after sampling noise disappears.The paper states that the expected leaderboard inflation grows significantly larger and does not vanish asymptotically.
- I.3.1 Extreme-value uplift: 56 Arena Score points is the reported bias for σtrue = 20 Arena Score, N = 50, and n = 3 000.Heterogeneous checkpoints have distinct underlying performance rather than only statistical fluctuations.
K Silent Model Deprecation: Additional Details
Many public models receive little Arena activity and are silently deprecated, with open-weight and open-source models affected more heavily than proprietary models. The analysis uses March 3–April 23, 2025 activity and task-specific simulation win rates.
- K Silent Model Deprecation: Additional Details: Top-10 sampling weights and higher weights for new models contribute to active sampling of some providers’ models and silent deprecation of others.Google, OpenAI, Anthropic, Amazon, Meta, and DeepSeek AI have between 3 and 10 actively sampled public models.
- K Silent Model Deprecation: Additional Details: 86.6% of open-weight models and 87.8% of open-source models are affected by silent deprecation, versus 2.4% of officially deprecated models that are open-weight.Among officially deprecated models, 30% are proprietary.
- Simulation setup: The simulation assigns task-specific pairwise win probabilities and uses separate tables for task-1 and task-2.Table 7 notes a 0.1 tie rate for A versus B, while Table 8 notes 0.2 tie rates for A versus B and A versus C.
M Overfitting Experiments: Additional Evaluations
Training on Arena battle data improves performance on Arena-specific evaluation, but has little to no effect on the non-Arena MMLU benchmark.
- All models achieve very similar accuracy on MMLU despite training with varying amounts of Arena data.This contrasts with the gains observed on ArenaHard prompts as Arena-mix data increases.
- Arena battle data boosts scores specific to Arena evaluation while providing little to no benefit on non-Arena benchmarks.