Source-linked AI summary

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Rose Kirk, Fangru Lin, Gabrielle Kaili-May Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yushi Yang, Yilun Zhao, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip H. S. Torr, Cozmin Ududec, Luc Rocher, Adam Mahdi

arXiv:2511.04703v1cs.CLcs.AI

TL;DR

LLM benchmarks must validly represent abstract phenomena because their scores support claims about capabilities, safety, and robustness. The paper systematically reviews 445 benchmarks using 29 expert reviewers and finds widespread weaknesses in operationalisation, task design, metrics, and interpretation. It responds with eight recommendations and a practical checklist for improving construct validity.

  • Problem

    Abstract LLM phenomena cannot be measured directly, so benchmarks need construct-valid proxies that support reliable claims about their targets.

  • Method

    The authors systematically reviewed 445 benchmarks from leading ML and NLP conferences using a detailed schema applied by 29 experts.

  • Results

    Nearly every reviewed benchmark had weaknesses in at least one area, including insufficient operationalisation, unrepresentative tasks, and rare statistical testing.

  • Takeaways & Limitations

    The paper proposes eight recommendations and an operational checklist for benchmark design and interpretation.

  • Takeaways & Limitations

    The review focuses on leading conference proceedings and may exclude impactful industry benchmarks without formal peer review or benchmarks from specialised venues.

Abstract

from arXiv · show

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as 'safety' and 'robustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks.

1 Introduction

LLM benchmarks operationalise abstract phenomena through tasks and metrics, so their value depends on construct validity: whether scores support claims about the intended phenomenon. This review examines weaknesses in those practices and proposes improved standards.

  • Benchmarks translate abstract phenomena into concrete tasks and metrics that act as measurable proxies for model capabilities.The benchmark’s value depends on whether that proxy represents the real-world phenomenon it intends to measure.
  • Construct validity is the degree to which a benchmark score provides evidence for claims about its target phenomenon.Low construct validity can make a high score irrelevant or misleading.
  • The review assesses construct-validity practices across 445 articles from leading machine-learning and natural-language-processing conferences.Experts coded phenomena, tasks, metrics, and claims using a detailed conceptual and methodological schema.
  • Nearly every reviewed article had weaknesses in at least one area, while key concepts were often poorly defined or operationalised.These weaknesses limit the reliability of the conclusions drawn from benchmark results.
  • The paper calls for improved reporting standards and releases an operational checklist of best-practice recommendations for establishing construct validity.The checklist is intended to support researchers and practitioners developing LLM benchmarks.

2 Background

Construct validity concerns whether tests measure the phenomena they intend to measure, especially when those phenomena cannot be directly verified. For LLMs, this issue is increasingly important as evaluations target broad abilities whose interpretation remains contested.

  • Construct validity evaluates whether an empirical test measures the phenomenon it intends to measure.The concept originated in psychological testing for phenomena that cannot be directly verified, such as personality.
  • Face, content, ecological, predictive, convergent, discriminant, and criterion validity assess different features of test design.These dimensions address representation of the phenomenon, task content, real-world relevance, future performance, and relationships with other measures.
  • Construct validity is central to benchmarking abstract LLM abilities such as reasoning.Standard benchmarks and narrowly defined tasks are becoming saturated as attention shifts toward general-purpose abilities.
  • Interpretations of evaluations remain contested, including whether results indicate intelligence or emergent abilities.These disagreements make assessment of construct validity increasingly important.

3 Methods

The study systematically reviews benchmark papers from major ML and NLP conference proceedings, combining model-assisted screening with manual expert review. A structured codebook captures validity-relevant features of phenomena, tasks, metrics, and claims.

  • Study design: The corpus contained 46,114 articles from ICML, ICLR, NeurIPS, ACL, NAACL, and EMNLP proceedings published between 2018 and 2024.The ACL range was limited by abstract availability.
  • Study design: Keyword screening identified 2,189 articles whose titles or abstracts mentioned benchmark and either LLM or language model.Most identified articles came from recent years.
  • Study design: Four inclusion criteria filtered papers for LLM capabilities, empirical benchmarks with reported performance, and compatibility with the review scope.Articles focused solely on technical aspects, opinions, reviews, and policy frameworks were excluded.
  • Study design: GPT-4o mini screened papers against the first three criteria with an F1 score of 84% on 50 human-labelled articles, reducing the set to 522.Twenty-nine expert reviewers then manually filtered the eligible papers, yielding 445 included articles.
  • Codebook and expert review: The codebook treated a benchmark as a task and metric used together to represent a phenomenon, alongside authors’ interpretation of results.Its items were derived from prior literature on face, predictive, content, ecological, convergent, and discriminant validity.

4 Results

The reviewed benchmarks span diverse phenomena, tasks, item sources, response formats, and metrics, but often rely on constructed or reused materials and provide limited statistical or validity evidence. These patterns motivate recommendations for more representative task design and stronger interpretation.

  • Review corpus: The review dataset covered 21 question items across 445 benchmark articles, annotated by 29 NLP and machine-learning experts.The annotations covered phenomenon definitions, task selection, metrics, and claims about measurement.
  • Phenomenon: Reasoning comprised 18.5% of reviewed phenomena, alignment 8.1%, and code generation 5.7%; 78.2% of articles defined their measured phenomenon.Among those definitions, 47.8% were contested rather than widely agreed upon.
  • Phenomenon: 61.2% of benchmarks defined their phenomenon as composite, compared with 36.5% defining it as a single unified whole.Composite phenomena integrate multiple sub-abilities, such as intent recognition, alignment, and structured output generation for agentic capabilities.
  • Task: Less than 10% of benchmarks used complete real-world tasks, while 40.7% used constructed tasks and 28.5% used them exclusively.Partially real-world and representative tasks appeared in 32.3% and 36.9% of benchmarks, respectively.
  • Task: Only 33.6% of benchmarks relied on a single task source; 43.3% handcrafted items, 42.6% reused existing benchmark data, and 31.2% generated data with LLMs.Human exams and other pre-existing sources appeared in 38.2% of benchmarks.
  • Task sampling: 12.3% of benchmarks used only readily accessible datasets, while another 27.0% incorporated convenience sampling in their strategy.The authors describe this as sampling from an inadequately controlled task space, which can limit validity.
  • Metric: Exact matching was used at least partially by 81.3% of benchmarks and exclusively by 40.7%, while only 16.0% used uncertainty estimates or statistical tests.Other metrics included soft matching, LLM-as-a-judge, and human ratings.
  • Claims: 53.4% of articles presented evidence for construct validity, with comparisons to similar benchmarks in 35.2% and human baselines in 32.4%.Comparisons to more realistic settings appeared in 31.2% of articles.

5 Recommendations

The review recommends treating construct validity as a lifecycle concern: define phenomena precisely, isolate them in tasks, sample representative items, and validate scoring choices. These practices address confounding, undersampling, and unclear operationalisations that weaken benchmark interpretation.

  • Define the phenomenon: Benchmarks should provide precise, scoped definitions of target phenomena and separately measure identifiable sub-components.78.2% provide definitions; 47.8% report no consensus on definitions, while 61.2% involve phenomena with sub-components.
  • Measure only the phenomenon: Benchmark tasks should control unrelated abilities, format constraints, and output parsing that can confound measurement of the intended phenomenon.21.1% of benchmarks require specific output formats that can challenge models independently of the target construct.
  • Construct a representative dataset for the task: Task datasets should use sampling strategies that represent the broader task space and include quality checks and known LLM sensitivities.Convenience sampling appeared in 27.0% of benchmarks, whereas 17.1% used at least one random or stratified sampling method.

5.5 Prepare for contamination

The recommendations emphasize contamination controls, statistically grounded comparisons, error analysis, and explicit validity rationales. The GSM8K example illustrates both useful practices and unresolved weaknesses in interpreting benchmark scores.

  • Prepare for contamination: Benchmark creators should test for contamination, maintain held-out items, and investigate exposure of source materials in training corpora.Contamination can affect validity through direct item exposure, memorisation of partial answers, or closely related information.
  • Use statistical methods to compare models: Only 16.0% of reviewed benchmarks conducted any statistical testing, despite recommendations to report sample sizes, uncertainty, rater characteristics, and rating variability.These practices support more robust comparisons, particularly for subjective human- or LLM-rated metrics.
  • Conduct an error analysis: Error analysis should characterize common failure modes, test for confounders, and identify scoring biases that affect construct validity.When failures reflect the target phenomenon, the resulting improvement directions can indicate higher validity; otherwise, validity may be reduced.
  • Justify construct validity: Authors should justify the chain from phenomenon definition through task, item, implementation, and validity claims, while discussing design trade-offs.Only 53.4% of reviewed benchmarks justified why they validly measured an important phenomenon.
  • GSM8K demonstration: GSM8K is generally valid for grade-school mathematics performance, but likely contamination, absent error analysis, and broad reasoning claims limit interpretation.The authors recommend clearer discussion of how its reading-comprehension and logical-reasoning requirements relate to reasoning broadly.

6 Discussion

The systematic review finds widespread weaknesses in how LLM benchmarks define phenomena, construct tasks, use statistics, and interpret validity. The authors respond with recommendations and a checklist intended for both new and existing benchmarks.

  • Findings: Across 445 reviewed benchmarks, abstract phenomena were often insufficiently operationalised, tasks frequently reused unadjusted data, and statistical testing was rare.About half discussed benchmark validity, but nearly every paper had at least one weakness.
  • Recommendations: The authors created recommendations covering phenomena, tasks, metrics, and interpretation to improve construct validity in future LLM benchmarks.The recommendations are framed as responses to the review’s identified gaps.
  • Operational checklist: The operational checklist supports construct-validity decisions throughout benchmark design and interpretation, while allowing researchers to document skipped items and trade-offs.It can also evaluate existing benchmarks or support adaptation to new domains and capabilities.
  • Limitations: The review primarily covers mainstream academic benchmarks from leading conferences and may omit industry releases and specialised domain benchmarks.Conference selection ensured a baseline of peer-reviewed quality but constrained coverage.
  • Limitations: Automated preliminary screening, language-use shifts, and limited reviewer allocation may have introduced false negatives, older-paper underrepresentation, or less robust reviews.GPT-4o mini screening was validated against human annotation but may still have produced undetected systematic errors.

7 Conclusion

Robust evaluation is needed as LLMs advance, yet the review finds construct-validity gaps across benchmarks. The paper proposes recommendations and a checklist, and calls for sustained attention to rigorous validation.

  • Conclusion: The review of 445 benchmarks identifies prevalent construct-validity gaps that undermine accurate measurement of targeted phenomena.The authors frame these gaps as a challenge for robust LLM evaluation.
  • Conclusion: Eight recommendations and a practical checklist are proposed for designing and interpreting LLM benchmarks.The proposed tools address shortcomings identified in the systematic review.
  • Conclusion: Measuring what matters requires sustained community effort toward more explicit and rigorous validation of evaluation methodologies.The conclusion characterizes this as a cultural shift toward prioritising construct validity.

Supplementary Material

The supplementary material presents a construct-validity checklist covering phenomenon definitions, task design, metrics, analysis, claims, and benchmark trade-offs.

  • Phenomenon: The checklist recommends precise phenomenon definitions, explicit scope boundaries, and separate measurement of sub-components.It also asks researchers to acknowledge excluded aspects when defining the target phenomenon.
  • Task and Dataset: Task recommendations emphasize representative and relevant items, controls for unrelated tasks, format-constraint analysis, and tests of known LLM sensitivities.The checklist also recommends held-out items and contamination testing.
  • Metrics: Evaluation guidance calls for justified sample sizes, uncertainty estimates, demographic-bias mitigation, and metrics that capture variability in subjective labels.It cautions against relying on single-point aggregation or exact matching for subjective labels.
  • Analysis: The checklist recommends analyzing failure modes, testing confounders, and discussing scoring biases revealed by error analysis.These analyses connect observed errors to intended constructs and non-targeted phenomena.
  • Claims and Trade-offs: Researchers should connect benchmark claims to operational definitions, real-world applications, related evaluations, model-improvement experiments, and construct-validity trade-offs.The supplementary material notes that adopting every recommendation may be difficult and encourages discussing implementation trade-offs.
  • Codebook: The review dataset records article, benchmark, phenomenon, task, and methodological information for systematic analysis.Its codebook includes article identifiers, contribution summaries, phenomenon categories, definitions, and task descriptions.

B.5 Results and Claims

The reviewed benchmarks often lacked comparisons to realistic settings, similar benchmarks, or human performance, limiting the evidential basis for interpreting their claims.

  • Results and Claims: 294 benchmarks made no comparison to results on other benchmarks of similar phenomena, whereas 160 did.
  • Results and Claims: 261 benchmarks made no comparisons whose similarities or differences were explained by theories; 144 did, and 27 reported no comparison category.
  • Results and Claims: 305 papers did not present a human baseline on the task, while 146 did.
  • Results and Claims: Only 23 benchmarks compared results with more realistic settings, while 308 made no such comparison and 119 were themselves realistic.

C.1 Keyword Search and LLM Filtering

The review used keyword search, LLM-assisted screening, and manual filtering to reduce an initial corpus to the benchmark articles included in the study.

  • Keyword Search: 2,189 articles matched the title-or-abstract keywords ‘benchmark’ and ‘LLM’ or ‘language model’ across six target conferences.
  • LLM Filtering: GPT-4o mini progressively screened the keyword-selected articles against inclusion and exclusion criteria before human review.
  • LLM Filtering: 80% precision and 89% recall were achieved when LLM filtering was validated against human labels for 50 randomly selected articles.The authors describe the filtering as effective but not perfect.
  • Manual Filtering: 522 articles proceeded to manual review, and 445 were included in the final study, corresponding to about 85% precision among manually reviewed articles.

D Inter-rater Agreement

Inter-rater agreement was assessed on 46 papers across 30 categorical codebook fields, revealing moderate overall consistency and lower agreement in interpretive fields.

  • Study Design: 46 papers were independently annotated by two reviewers using 30 categorical items covering phenomena, tasks, metrics, and validity claims.
  • Agreement Measures: Brennan–Prediger Kappa was selected alongside percent agreement because the codebook contained binary, multi-class, and multi-label questions with imbalanced labels.
  • Agreement Measures: For multi-label fields, Jaccard similarity was averaged for percent agreement and thresholded at 0.3 before binary BPK computation.
  • Results: 68.1% mean percent agreement and 0.524 mean BPK indicated moderate consistency across the 30 annotated fields.
  • Results: Interpretive or compositional fields such as task_ecology and dataset_sampling_method showed lower consistency than structured and objective fields.
Loading 2511.04703v1…