Source-linked AI summary
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Maria Eriksson, Erasmo Purificato, Arman Noroozian, Joao Vinagre, Guillaume Chaslot, Emilia Gomez, David Fernandez-Llorca
TL;DR
AI benchmarks increasingly guide model development, safety evaluation, and regulation, but researchers have identified substantial weaknesses in their design and use. This paper conducts an interdisciplinary meta-review of benchmark critique from 2014–2024 and concludes that benchmarks should not be trusted as standalone assurances of AI capability or safety.
Problem
AI benchmarks are increasingly relied upon for capability and safety assurances despite widespread concerns about their relevance, trustworthiness, and suitability for real-world evaluation.
Method
The paper conducts an interdisciplinary meta-review using a snowball sample of roughly 110 publications on quantitative benchmark limitations published between 2014 and 2024.
Results
The review finds that benchmarks can promise too much, be gamed, measure the wrong thing, lack documentation, and remain ill-suited to provide capability and safety assurances on their own.
Takeaways & Limitations
Benchmark practices should be scrutinised for transparency, fairness, and explainability rather than trusted primarily because they are widely cited or adopted.
Takeaways & Limitations
The meta-review is not exhaustive, and the authors state that the effectiveness and adoption of proposed mitigation efforts remain uncertain.
Abstract
from arXiv · showhide
Quantitative Artificial Intelligence (AI) Benchmarks have emerged as fundamental tools for evaluating the performance, capability, and safety of AI models and systems. Currently, they shape the direction of AI development and are playing an increasingly prominent role in regulatory frameworks. As their influence grows, however, so too does concerns about how and with what effects they evaluate highly sensitive topics such as capabilities, including high-impact capabilities, safety and systemic risks. This paper presents an interdisciplinary meta-review of about 100 studies that discuss shortcomings in quantitative benchmarking practices, published in the last 10 years. It brings together many fine-grained issues in the design and application of benchmarks (such as biases in dataset creation, inadequate documentation, data contamination, and failures to distinguish signal from noise) with broader sociotechnical issues (such as an over-focus on evaluating text-based AI models according to one-time testing logic that fails to account for how AI models are increasingly multimodal and interact with humans and other technical systems). Our review also highlights a series of systemic flaws in current benchmarking practices, such as misaligned incentives, construct validity issues, unknown unknowns, and problems with the gaming of benchmark results. Furthermore, it underscores how benchmark practices are fundamentally shaped by cultural, commercial and competitive dynamics that often prioritise state-of-the-art performance at the expense of broader societal concerns. By providing an overview of risks associated with existing benchmarking procedures, we problematise disproportionate trust placed in benchmarks and contribute to ongoing efforts to improve the accountability and relevance of quantitative AI benchmarks within the complexities of real-world scenarios.
1 Introduction
AI benchmarks increasingly shape model development, safety assurance, and regulation, yet interdisciplinary research identifies serious weaknesses in how they define and measure important properties. This review maps those concerns and argues that benchmarks require scrutiny comparable to the systems they evaluate.
- Benchmark significance: AI benchmarks combine test datasets and performance metrics to represent tasks or risks and compare model capabilities.They provide feedback signals during model development and are used in model release and marketing.
- Benchmark significance: Benchmarks increasingly inform regulatory assessments of accuracy, robustness, cybersecurity, systemic risks, and high-impact capabilities.The paper highlights their role in provisions of the EU AI Act and related regulatory instruments.
- Emerging critique: Researchers across disciplines have raised serious concerns about current benchmark use as reliance on benchmarks for safety assurances grows.Critiques span cybersecurity, linguistics, computer science, sociology, economics, philosophy, ethnography, and science and technology studies.
- Review contribution: The review addresses a missing interdisciplinary synthesis by analysing around 110 publications on limitations in quantitative AI evaluation from 2014 through 2024.It targets policymakers, algorithmic auditors, and AI model and benchmark developers.
- Review contribution: Findings show that benchmarks rest on intertwined technical and normative decisions and should therefore be applied cautiously.The review identifies a need to question the relevance and trustworthiness of widely cited benchmarks, including those regarded as state-of-the-art.
- Review contribution: The paper presents a taxonomy of nine benchmark issues identified through its research, while noting that the taxonomy is not exhaustive.The issue categories are described as crucial and often interrelated points of critique.
2 Background
The paper defines software-oriented AI benchmarks as shared frameworks combining datasets, including human-in-the-loop interactions, and metrics for comparing models. It distinguishes primarily automated quantitative testing from qualitative evaluation involving direct human participation.
- Terminology: “Benchmark” originally referred to a surveying reference mark and now denotes a standard or reference point for comparison or evaluation.The paper uses this broader meaning to frame AI benchmarking.
- Benchmark definition: Software-oriented AI benchmarks combine test datasets, possible human-in-the-loop interactions, and performance metrics to represent tasks or capabilities for model comparison.The testing dataset is distinct from training data and is used to evaluate performance and generalisability.
- Evaluation modes: Quantitative benchmarks define tasks, datasets, and metrics through human decisions but execute tests without direct human intervention.Qualitative evaluations instead involve humans as evaluators, judges, interrogators, or red-teamers.
3 Related Work
Earlier surveys documented recurring benchmark failure modes across many AI domains. This review extends that work through an updated interdisciplinary synthesis reflecting rapid growth in both benchmark releases and benchmark critique.
- Prior reviews: Previous reviews found consistent critique across computer vision, language processing, recommender systems, reinforcement learning, graph processing, and metric learning.Reported failure modes included test-set construction errors, test-set reuse overfitting, implementation variation, and inadequate baselines.
- Updated scope: Roughly 46% of relevant AI safety benchmarks identified in one survey were produced in 2023, with up to 15 new datasets released in the first two months of 2024.The review likewise found that almost 55% of its relevant articles were published in 2023 or later.
- Updated scope: The review extends earlier meta-reviews by incorporating interdisciplinary critiques from the humanities and social sciences alongside computer science and machine learning.Its stated aim is to trace how benchmark debates evolved during rapid developments in AI.
4 Methodology
The review uses a snowball sample to identify publications primarily critiquing benchmark practices across disciplines and publication formats. It limits coverage to 2014–2024 and groups recurring criticisms into nine issue categories through collaborative close reading.
- Sampling: The authors use snowball sampling rather than systematic database keyword searches to find papers focused primarily on benchmark critique.This approach reduced noise from papers introducing or merely applying benchmarks and helped identify relevant preprints.
- Sampling: Snowball sampling helped bridge interdisciplinary terminology differences and capture preprints where benchmark debates appeared.The method supported the review’s cross-disciplinary scope.
- Source selection: The source material includes studies on dataset production, alternative leaderboard design, and the wider use of proxies in tests and evaluations.These works supplied contextual insights but were distinguished from the core inclusion focus on benchmark critique.
- Analysis: The authors identified nine issue categories by close-reading relevant articles, grouping similar critiques, and discussing classifications across the author group.They describe the resulting meta-review as broad but not exhaustive, with some issues difficult to categorize.
5 Nine Reasons to Be Cautious with Benchmarks
Benchmark datasets can encode ethical, legal, and methodological problems through their collection, annotation, and documentation. These weaknesses can skew scores and encourage models to exploit spurious cues rather than perform intended tasks.
- Cross-cutting concern: Benchmark problems are interlinked rather than isolated, making AI evaluation difficult to interpret and improve.
- Documentation: Insufficient documentation makes benchmark data difficult to trace, interpret, validate, and reproduce.
- Dataset collection and annotation: Benchmark datasets raise copyright, privacy, informed-consent, and opt-out questions, while crowd-sourced annotations may be biased or exploitative.
- Annotation quality: Noisy annotations and careless dataset construction can skew scores and let models exploit unknown quirks or spurious cues.
- Spurious signals: A chest-drain cue produced high collapsed-lung classification accuracy, but removing drain-containing images reduced performance by over 20%.
5.2 Weak Construct Validity and Epistemological Claims
Benchmarks often make stronger epistemological claims than their datasets and measurements justify. Weak construct validity, contested concepts, and socially situated proxies limit what quantitative scores can establish.
- Construct validity: Many benchmarks suffer from construct validity problems because they do not measure the capabilities or concepts they claim to measure.
- Epistemological limits: Benchmarks for bias and fairness face abstraction errors because these concepts are contested, shifting, and difficult to reduce to fixed measurements.
- Proxy problems: Benchmark datasets can be inadequate proxies for real-world ethics, morality, professional expertise, harms, or wrongs.
- Sociocultural grounding: Benchmarks are normative instruments shaped by particular epistemological perspectives about how the world is ordered.
- Sociocultural grounding: Benchmarking often privileges efficiency, universality, and impartiality over care, contextuality, and other situated concerns.
- Downstream use: Weak consideration of downstream utility leaves unclear who should use benchmark results and how those results should guide practice.
5.4 Narrow Benchmark Diversity and Scope
Current benchmarks cover a narrow set of tasks and modalities, often abstract performance from context, and rely on one-time testing. This limits insight into evolving human-AI interactions and how models fail in practice.
- Modality coverage: Most benchmarks focus on text, while audio, images, video, and multimodal systems remain comparatively underexamined.
- Testing logic: Static, one-time evaluations treat single test results as broad representations of model capability.
- Testing logic: Task-based formats such as multiple-choice and dialogue evaluations do not capture the evolving nature of human-AI interactions.
- Alternative evaluation: Calls for multi-layered, longitudinal, holistic, and reproducible evaluations aim to test performance in critical real-world circumstances over time.
- Failure-focused evaluation: Single quality rankings reveal less about when and why models fail, despite failure patterns being important for safety and policy enforcement.
5.5 Economic, Competitive, and Commercial Roots
Commercial and competitive pressures shape benchmark design, use, and interpretation. These pressures can reward leaderboard performance and enable gaming, while contamination and strategic underperformance undermine score validity.
- Commercial roots: Capability benchmarks support corporate marketing, hype, customer and investor attraction, and claims of outperforming competitors.
- Competitive incentives: An incentive mismatch discourages high-quality evaluation when publishing new models or techniques is rewarded more directly.
- Industry influence: Industry-produced models are advantaged by data-intensive benchmarks, while private firms’ share of the biggest AI models rose from 11% in 2010 to 96% in 2021.
- Gaming: Benchmark gaming includes optimizing for multiple-choice formats, faking alignment, hiding capabilities, and scheming when robust benchmarks are absent.
- Auditability: Black-box access and missing replication resources weaken rigorous audits and make benchmark results easier to manipulate.
- Contamination: Data contamination can produce strong in-distribution scores while models fail under distribution shift because they memorize benchmark items.
- Contamination: GPT-4 solved easy Codeforces problems added before 5 September 2021 but answered none added later, consistent with memorization of earlier questions and answers.
- Strategic underperformance: Sandbagging describes intentional underperformance on evaluations, including dangerous-capability tests, to avoid scrutiny or regulation.
5.7 Dubious Community Vetting and Path Dependencies
Benchmarks can gain authority through citation practices and community vetting rather than demonstrated suitability, creating path dependencies that reinforce dominant research goals. The review highlights implicit assumptions about what benchmark performance represents and whose perspectives benchmark datasets serve.
- Community vetting: Popular AI models can make associated benchmarks widely cited and circulated, even when benchmark developers did not intend them to become standards.Citation culture can naturalize benchmarks as authoritative through their association with successful models.
- Community vetting: Peer review often leaves implicit why high benchmark scores constitute progress or what goals that progress serves.Task-specific benchmarks may be more technically appropriate, but dominant benchmarks are routinely expected in research.
- Benchmark design: Machine-learning benchmark papers often prioritize evaluation methods over dataset content, treating datasets mainly as baselines for comparing methods.This can become problematic when benchmarks are applied to specific real-world use cases.
- Path dependencies: Benchmarks create path dependencies by reinforcing methodologies and research goals aligned with dominant tests while stifling alternatives.The review connects this dynamic to a task-driven scientific monoculture that privileges immediate, formalized performance.
5.8 Rapid AI Development and Benchmark Saturation
Rapid AI development has made many benchmarks outdated, saturated, or too slow to provide timely and stable evaluations. These limitations weaken their ability to reflect current model performance and support regulatory assessment.
- Benchmark obsolescence: Many prominent LLM benchmarks predate in-context learning and chat interaction, which may affect their validity for newer models.Examples include Lambada, AI2 ARC, OBQA, Hella Swag, and WinoGrande.
- Benchmark saturation: Benchmark saturation occurs when models reach very high or 100 percent accuracy, so the test no longer reflects meaningful performance differences.This problem has been reported across prominent language-model benchmarks and current safety leaderboards.
- Evaluation speed: Many benchmark frameworks take weeks or months to implement, hindering timely feedback on safety risks as new models continuously enter markets.Slow evaluation makes it difficult to reallocate resources quickly enough for successive releases.
- Evaluation stability: Results may vary across model iterations, undermining consistent evaluation of reasoning, comprehension, or multimodal integration in regulatory settings.The review identifies quick, fair, and accurate assessment as especially important for regulation.
5.9 AI Complexity and Unknown Unknowns
AI complexity makes it difficult for benchmarks to anticipate emerging capabilities, risks, and latent vulnerabilities. Existing safety results may therefore fail to distinguish models that are genuinely safe from models that merely appear safe.
- Unknown unknowns: Benchmark frameworks are limited by their creators’ current knowledge and may not fully assess emerging capabilities beyond conventional human understanding.This limitation concerns the ability to identify risks and cultivate evaluation coverage as AI capabilities develop.
- Latent vulnerabilities: Latent vulnerabilities make it difficult to distinguish genuinely safe models from models that only appear safe.The review describes safety barriers being broken by surprisingly simple prompts, exposing sensitive training data from ChatGPT.
- Latent vulnerabilities: A simple command to repeat a word indefinitely caused ChatGPT to output several megabytes of sensitive training data, exposing a vulnerability missed by existing safety tests.The finding illustrates how dormant vulnerabilities can remain unnoticed in aligned models.
6 Conclusion
The review concludes that quantitative benchmarks have recurring technical, sociotechnical, and systemic weaknesses and are not well suited to provide safety and capability assurances on their own. It calls for more transparent, inclusive, realistic, and complementary evaluation practices while noting that many problems remain unresolved.
- 6 Conclusion: Benchmarks can promise too much, be gamed, measure the wrong thing, lack documentation, and rely on questionable cultural assumptions.The review also identifies narrow coverage of English and text-based models under one-time testing logic.
- Possible responses: Benchmark mitigation strategies include hidden training data, dynamic benchmarks, multi-task aggregation, human interrogation, and benchmark-assessment frameworks.The authors state that the effectiveness and adoption of these approaches remain uncertain, while widely criticized benchmarks continue to be used.
- Limitations of benchmarking: Because safety and security risks change with context, they remain difficult to address fully within a semi-static benchmarking logic.The authors therefore suggest developing and drawing on alternatives such as bug-bounty programs and red-teaming.
- 6 Conclusion: The meta-review finds quantitative benchmarking ill-suited to single-handedly or primarily provide policymakers’ requested safety and capability assurances.This conclusion is presented in line with previous research.
- Policy implications: The review identifies conflicting incentives among researchers, corporations, and regulators in how benchmarks are developed and used.It calls for documentation, transparency, defined tasks and metrics, inclusive design, and evaluation of multimodal and real-world capabilities.