Source-linked AI summary
The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
Jason Branford, Angelie Kraft
TL;DR
AI benchmarks are often treated as technical measures, although they shape power, recognition, and research priorities. Using Iris Marion Young’s theories, the paper argues that benchmarking perpetuates structural injustice and undermines AI research epistemically and ethically.
Problem
AI benchmarking is treated as technical measurement despite its wider effects on power, research priorities, and marginalised communities.
Method
The paper analyses benchmarking culture through Iris Marion Young’s theories of oppression and structural injustice.
Results
Benchmarking perpetuates exploitation, marginalisation, powerlessness, and cultural imperialism while structurally undermining AI research.
Takeaways & Limitations
The paper calls for democratic, localised, context-specific, use-specific, and recurring evaluation with meaningful participation by diverse stakeholders.
Takeaways & Limitations
The paper does not claim that all benchmark researchers and developers are blameworthy, locating the harms in collective, often well-intentioned practices.
Abstract
from arXiv · showhide
Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of leaderboards that reward state-of-the-art performance with prestige, citations, trust, and institutional influence. As the costs of developing competitive AI systems rise, these rewards increasingly concentrate among powerful, industry-funded labs. This paper situates these concerns within Iris Marion Young's theories of oppression and structural injustice. It argues that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Young's "faces of oppression". Benchmarking culture is further framed as a source of structural injustice, as these harms emerge from normalised, individually defensible practices and network effects, even without explicit wrongdoing. By reinforcing existing power structures and narrowing possible research trajectories, benchmarking may in fact prevent the field from advancing in epistemically robust and socially beneficial ways.
1 Introduction
AI benchmarks do more than measure system performance: prominent benchmarks help define how AI capabilities are understood and communicated. The paper argues that benchmarking culture may reproduce oppression and injustice affecting smaller labs, marginalised researchers, and marginalised users, thereby hindering the field’s potential.
- 1 Introduction: Prominent benchmarks turn limited technical performance into public evidence about AI capabilities, shaping how systems are understood and the ends to which they are put.Their influence is especially pronounced when they reach audiences beyond specified academic sub-communities, are widely used for communication, and appear on leaderboards.
- 1 Introduction: Benchmarking culture may contribute to structures of oppression and injustice affecting smaller labs, marginalised researcher communities, and marginalised user groups.The paper treats these effects as occurring on both the research and user sides of AI.
- 1 Introduction: Drawing on Iris Marion Young, the paper examines the economy and epistemology of AI benchmarking and argues that benchmarking may hinder the field’s potential.The analysis assesses how current benchmarking culture relates to oppression and injustice.
2 The Influence and Troubles of Benchmarks
AI benchmarks combine datasets and metrics to rank systems, but their validity, representativeness, and governance are contested. Because benchmark scores circulate among diverse stakeholders and shape development, deployment, trust, and public claims, these shortcomings have broader social consequences than technical measurement error alone.
- Benchmark foundations: Benchmarks combine datasets, ground-truth labels, and metrics to calculate scores that compare and rank AI systems.Their labels distinguish true from false or desired from undesired responses.
- Stakeholders and uses: Benchmarks serve researchers, engineers, decision-makers, operators, regulators, journalists, data authors, and annotators across different evaluation purposes and contexts.Journalists draw on and thereby perpetuate benchmark leaderboards and scores when reporting on AI progress.
- Validity and bias: More than half of 445 reviewed LLM benchmarks relied on contested or undefined concepts, supporting longstanding concerns about weak construct validation.The review provides empirical support for critiques that computer scientists often lack clear concept definitions and evidence-based validation.
- Validity and bias: Biased or unrepresentative benchmark data can miscalibrate evaluation, reward biased outputs, and feed algorithmic bias into subsequent systems.Once adopted, benchmark scores can harden across domains, so methodological shortcomings cannot be assessed separately from their social context.
- Competitive dynamics: Benchmark winners gain praise, recognition, downloads, citations, and public and institutional trust, creating incentives that structure the community’s innovation efforts.Benchmarks also become de facto standards that influence which models stakeholders develop further or deploy.
- Governance and power: Data contamination, benchmark gaming, and secretive benchmark partnerships undermine trustworthy evaluation and turn benchmarks into tools of power for resourceful institutions.The passages identify the FrontierMath partnership and admitted manipulation of Llama 4 results as warning signals, while noting that widespread intentional misconduct is difficult to prove.
3 Benchmarking and the Faces of Oppression
AI benchmarking practices can perpetuate oppression by channeling collective labour toward standards controlled by powerful institutions, marginalising less-resourced researchers and communities, and rendering dominant perspectives authoritative. These harms arise through normalised, often well-intentioned collective decisions rather than requiring blameworthy complicity by individual benchmark creators or users.
- Exploitation: Benchmarks coordinate labour, attention, prestige, and investment, directing researchers toward public scores whose terms and rewards are disproportionately controlled by powerful institutions.Benchmark designers can shape research purposes, success criteria, and the conversion of gains into institutional and economic advantages.
- Exploitation: Evaluative rules channel collective energy toward standards that many must follow but few can contest, entrenching dependence on powerful institutions and limiting disagreement over AI’s purposes.The paper frames this as exploitation through structural relations, not necessarily through individually blameworthy conduct.
- Cultural imperialism: Demographic and geographic biases, opaque reporting, and elite institutional control make benchmark authority reflect select perspectives while shaping which research counts as progress.Identified biases include preferences toward Western, Christian, and male subjects, alongside limited transparency about annotators.
- Marginalisation: Leaderboards increasingly presuppose large-scale compute, proprietary data, and extensive engineering labour, often measuring resource access more than ingenuity, scientific value, or social utility.This dynamic pushes researchers and entire communities from less-resourced institutions and regions toward the periphery.
- Marginalisation: English-specific benchmarks limit visibility and praise for work on low-resource languages, while marginalised communities’ own standards of success remain excluded from prevailing evaluation norms.Inclusion risks serving dominant metrics, market reach, or claims of generality rather than community self-determination.
- Powerlessness: Opaque evaluation infrastructure and leaderboard rankings deprive the public of independent judgment, encouraging users to treat benchmark metrics as reliable evidence of what is best.Users can become epistemically deferential to AI models as benchmarks dominate diagnoses of accuracy, robustness, and security.
4 Structural Injustice and the Science of AI
AI benchmarking constitutes structural injustice because normalised, collectively upheld practices can produce systematic harms without ill intent, while influential actors possess unequal power to preserve or change the system. These structures also weaken AI’s epistemic foundations by excluding less dominant groups and maintaining a closed evaluation cycle.
- Structural injustice: Benchmarking-related oppression arises from a culture of evaluation established over time and upheld through collective practices, rather than only from isolated harms or individual wrongdoing.These harms are enabled by benchmarking’s stable position within the AI ecosystem.
- Structural injustice: Normalised, seemingly unobjectionable actions by individuals can generate structural injustice through network effects, even when people act acceptably and without ill intent.This corresponds to McKeown’s account of pure structural injustice.
- Structural injustice: Influential benchmarking actors hold unequal power over the system and have the capacity to continue leveraging it for their benefit, change it, or mitigate its harms.The diversity of benchmarks and their uses makes the landscape difficult to categorise as a whole.
- Epistemic consequences: AI evaluation systems exclude certain groups from a significant epistemic process, inadequately representing less dominant researchers, developers, and consumers.The paper presents this exclusion as ethically and epistemically harmful because it produces insights about AI progress and appropriateness that are less relevant to their preferences.
- Epistemic consequences: AI research is trapped in a closed evaluation cycle in which those who set evaluation methods build the evaluated systems and benefit from their uptake.The proposed response is third-order change: revising the epistemic tools used to judge and tune AI systems.
5 Implications for the Future of AI Research
The paper calls for AI evaluation practices that acknowledge benchmarking’s structural injustices while pursuing more valid, reliable, transparent, participatory, and accountable forms of assessment. It emphasizes that evaluation must remain explicit about value judgments, situated knowledge, and the power relations shaping who evaluates and benefits.
- Structural injustice: AI benchmarking shapes labour, recognition, authority, and research direction while perpetuating exploitation, marginalisation, powerlessness, and cultural imperialism.The paper characterizes these harms as forms of structural injustice and argues they are avoidable.
- Evaluation science: Evaluation science should improve benchmarking’s statistical foundations, measurement instruments, documentation, validity, reliability, and transparency without focusing only on individual symptoms.The paper welcomes current efforts toward more rigorous evaluation while warning that they may overlook broader structural problems.
- Situated evaluation: Because benchmarking is reductive and prone to capture, better evaluation should recognize situated knowledge and make its value-laden purposes explicit.The paper presents these practices as ways to mitigate persistent deficiencies rather than eliminate them.
- Participatory evaluation: Better AI evaluation requires a wider participatory ecosystem of imagination, debate, and audit while accepting that judgments of what is right or wrong are local and temporary.The paper frames locality and temporality as conditions of evaluation rather than defects to be fully overcome.
- Participatory evaluation: Red-teaming and open voting platforms such as Chatbot Arena broaden participation, but they may be falsely marketed as democratic while powerful players harvest free labour.These approaches illustrate both collective evaluation and its vulnerability to power asymmetries.
- Accountability and oversight: The AI community must hold powerful actors accountable for capturing, exploiting, and gaming evaluations, while regulators should require true democratic oversight of AI tools.The paper presents accountability and oversight as necessary to mitigate avoidable structural injustices in benchmarking.