Source-linked AI summary

AI and the Everything in the Whole Wide World Benchmark

Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton, Alex Hanna

arXiv:2111.15366v1cs.LGcs.AIcs.PF

TL;DR

This position paper examines whether influential AI benchmarks can support claims about general capabilities, focusing on their specificity, finiteness, and contextuality. It analyzes benchmarking practice and construct validity, concluding that no dataset or current benchmarking method meaningfully measures general capabilities and that benchmarks should instead be contextualized and appropriately scoped.

  • Problem

    Influential benchmarks are often framed as broad measures of progress toward general AI abilities, despite being specific, finite, and contextual.

  • Method

    The paper defines benchmarks and construct validity, then examines benchmarking practice and benchmark examples to analyze how their design relates to claims about general capabilities.

  • Results

    No dataset can capture the full complexity of existence, and current benchmarking methods do not meaningfully measure general capabilities.

  • Takeaways & Limitations

    Benchmarks are more useful when reframed and contextualized as tools for understanding how systems work and where they do not.

  • Takeaways & Limitations

    Benchmark coverage is constrained by closed-world, temporally bounded, and geographically uneven datasets, while increasing size can trade off against annotation quality.

Abstract

from arXiv · show

There is a tendency across different subfields in AI to valorize a small collection of influential benchmarks. These benchmarks operate as stand-ins for a range of anointed common problems that are frequently framed as foundational milestones on the path towards flexible and generalizable AI systems. State-of-the-art performance on these benchmarks is widely understood as indicative of progress towards these long-term goals. In this position paper, we explore the limits of such benchmarks in order to reveal the construct validity issues in their framing as the functionally "general" broad measures of progress they are set up to be.

1 Introduction

The paper compares inflated claims about general AI benchmarks to Grover’s absurd museum of “everything,” arguing that finite, specific, contextual datasets cannot represent general ability. It examines how benchmarks such as GLUE and ImageNet became elevated beyond their original purposes, while warning against framing them as general measures.

  • Grover’s museum uses finite examples to claim coverage of “everything in the whole wide world,” exposing the absurdity of representing all categories through selected objects.Its rooms range from arbitrary categories such as “Things You Find On a Wall” to vague ones such as “The Tall Hall,” ending with an “Everything Else” door opening onto the outside world.
  • The paper argues that AI benchmarks make a similar mistake when vague abilities such as “visual understanding” or “language understanding” are represented by finite, specific, contextual datasets.
  • GLUE and ImageNet are often treated as definitions of essential common tasks, allowing claims about them to extend beyond the datasets’ initial designs and ambitions.
  • The paper does not deny benchmark utility; it investigates how the historical Common Task Framework evolved into benchmarks rhetorically presented as measuring general capabilities.

2 Background

The background distinguishes broad aspirations for flexible AI from historically practical benchmarking, then introduces construct validity as the question of whether evaluations appropriately characterize real-world behavior and research goals. Benchmarks combine datasets, tasks, and metrics into shared comparison frameworks, but their validity depends on how those components represent the intended claim.

  • 2.1 Striving for Generality: “Generality” commonly refers to AI systems that demonstrate competence across a wide range of tasks and settings, though this paper focuses on generality claims for particular skills such as vision or language.
  • 2.1 Striving for Generality: General-purpose representations aim to generalize with minimal fine-tuning to tasks for which they were not specifically developed.
  • 2.2 A Brief History of Benchmarking Practice in AI: Benchmarking has historically communicated application-based performance, whereas aspirations to measure general cognitive capabilities arise from broader scientific and philosophical views of intelligence.
  • 2.2 A Brief History of Benchmarking Practice in AI: A benchmark combines datasets, specific tasks, and a metric into a shared framework for comparing methods, with favorable scores defining state-of-the-art performance on the specified task.
  • 2.2 A Brief History of Benchmarking Practice in AI: Benchmarking developed from practical computer-selection and task-specific software tests toward common-task evaluation frameworks for computational linguistic and vision tasks.
  • 2.3 Construct Validity: Construct validity concerns how well an experimental setting, including its dataset, metrics, and methodology, relates to the research claim and intended real-world behavior.

3 Attempting to Benchmark General Capabilities

The paper examines ImageNet and GLUE as benchmarks whose presentation and community use embed claims of generality beyond their bounded construction. It argues that extending task-specific benchmark practices to abstract general capabilities exceeds what these datasets can validly establish.

  • Recent benchmarks are presented and adopted as if they embed generality, even when their creators do not explicitly claim to measure general intelligence.
  • The paper argues that data-defined benchmarks cannot adequately embody general-purpose object recognition, general language understanding, or domain-independent reasoning.
  • The analysis focuses on ImageNet and GLUE as dataset-based benchmarks for general capabilities, excluding substantial discussion of gameplay demonstrations.
  • ImageNet: ImageNet was described as comprehensive and diverse coverage of the image world, and increased ImageNet performance has been referenced as progress toward general-purpose AI.
  • GLUE and SuperGLUE: GLUE and SuperGLUE are framed as general-purpose language-understanding evaluations intended to cover varied training volumes, genres, and task formulations.

4 Limits of Benchmarking General Capabilities

The paper argues that benchmarks framed as measures of general capabilities are limited, subjective constructions whose tasks and datasets do not systematically represent broad abilities. Examples from ImageNet, GLUE, and SuperGLUE show arbitrary task selection, conflated competencies, restricted coverage, and culturally situated data.

  • 4.1 Limited Task Design: “General” benchmarks are not systematic abstractions of general functions but collections of inconsistently selected tasks and categories.The paper compares their construction to the arbitrary rooms in Grover’s museum.
  • 4.1.1 Arbitrarily Selected Tasks and Collections: ImageNet combines specific and high-level classes inherited largely from WordNet, including offensive categories and categories selected for consistency with PASCAL.Its taxonomy spans dog breeds, “New Zealand beach,” and broad WordNet subtrees.
  • 4.1.1 Arbitrarily Selected Tasks and Collections: GLUE’s tasks came from an informal survey and practical filtering, without systematically covering linguistic skills such as pragmatics or negation.The resulting collection reflects problems perceived as interesting by NLP researchers at the time.
  • 4.1.2 Critical Misunderstandings of Domain Knowledge and Application Problem Space: GLUE and SuperGLUE conflate linguistic competence with open-ended world knowledge and commonsense reasoning, encouraging overly broad interpretations of performance.The paper distinguishes reusable linguistic knowledge from open-ended world knowledge.
  • 4.1.2 Critical Misunderstandings of Domain Knowledge and Application Problem Space: Text-only tasks such as GLUE cannot thoroughly test world knowledge, commonsense reasoning, interlocutor modeling, or physical and social grounding.The paper notes that benchmark success may reflect manipulation of linguistic form rather than evidence of understanding.
  • 4.2 Limited Scope: ImageNet and GLUE are closed, localized, and subjective datasets whose limited coverage and annotation constraints resist interpretation as general measures.ImageNet is temporally bounded, underrepresents non-Western contexts, and faces trade-offs between dataset size and annotation quality.
  • 4.2.1 Limited Scope: GLUE’s task formats are narrow compared with human linguistic activity, prompting SuperGLUE to add question answering and other formats.GLUE consists mainly of sentence and sentence-pair classification tasks.
  • 4.2.2 Benchmark Subjectivity: ImageNet’s geographic sourcing is highly concentrated in Western countries, and Hindi-query construction would produce markedly different visual representations.The passage reports 45% of images sourced from the US and only 1% from China and 1.2% from India.

4.3 Inappropriate Community Use

The paper argues that community use of general benchmarks can redirect research toward leaderboard performance while obscuring contextual, subgroup, and real-world limitations.

  • Overstating benchmark generality can make datasets field-wide targets, encouraging algorithmic improvement that mismatches real-world or more relevant problems.
  • SOTA chasing emphasizes empirical, incremental comparisons over hypothesis-based scientific inquiry and can encourage metric manipulation and short-term goals.
  • Dataset viewpoints reflect their authors and annotators, so unexamined construction can over-represent hegemonic perspectives.
  • Aggregate benchmark scores can make improvements across datasets appear comparable even when their practical meanings differ substantially.
  • Competitions and leaderboards can redirect research toward selected topics and dominant approaches, as illustrated by FERET, Netflix Prize, and chess.
  • Marketing claims can further distort benchmark significance by presenting benchmark performance as a reliable marker of deployment achievement, potentially hiding subgroup disparities.

5 Alternative Roles for Benchmarking and Alternative Evaluation Methods

The paper rejects expanding supposedly general benchmarks as the solution and instead recommends contextualized task evaluation alongside alternative methods for broader objectives.

  • Treating benchmarks as independent of context, scope, and specificity is itself a false premise for machine learning evaluation.
  • Benchmarks should be developed, presented, and understood as evaluations of concrete, well-scoped, contextualized tasks.
  • For broader objectives, the paper proposes exploring testsuites, audits, adversarial testing, system-output analysis, ablations, and model-property analysis.

6 Conclusion

The conclusion argues that no dataset can represent the full complexity of the world or meaningfully measure general capability. Benchmarks remain useful when reframed as contextual tools for understanding system behavior.

  • Open-world, universal, and neutral datasets do not exist, so current benchmarking cannot meaningfully measure general capabilities.
  • Benchmark effectiveness depends on helping researchers understand how systems work and fail, not on making arbitrary claims of generality.

A Appendix A: Details of Alternative Evaluation Methods

The appendix reframes benchmarking through surveying: measurements should describe a changing landscape rather than merely rank state-of-the-art systems. It surveys alternative evaluation methods for that broader purpose.

  • A benchmark’s surveying-inspired meaning concerns understanding the shape of a landscape and how it changes, rather than measuring how far anyone has advanced.
  • Surveying requires relating measurements to derived results and accounting for external factors that influence those measurements.
  • Figure 1 depicts a physical benchmark associated with surveying, supporting the appendix’s metaphor for landscape-oriented evaluation.
  • The appendix reviews testsuites, audits, adversarial testing, system-output analysis, ablation testing, and analysis of model properties.

A.1 Testsuites, Audits and Adversarial Testing

Testsuites and audits deliberately map out categories of test items rather than relying on sampled benchmark distributions, while adversarial testing probes competence boundaries with minimally contrasting examples.

  • Testsuites and audits design test sets to map a space of item types and evaluate how extensively systems handle them.Typical benchmark test data instead reflects the frequency distribution of item types in an underlying dataset.
  • Audit-like evaluation balances sensitive categories to measure differential performance across those categories.
  • Adversarial testing explores competence boundaries using minimally contrasting pairs where a system succeeds on one example and fails on the other.

A.2 System Output Analysis

System output analysis examines errors, subgroup disparities, and counterfactual responses in detail, producing findings that are rich but unsuitable for quick cross-system comparison.

  • System output analysis includes error analysis, disaggregated analysis, and counterfactual analysis.
  • Error analysis: Error analysis inspects inputs and outputs mechanically or manually to identify reliably labeled, confused, or otherwise recurring error patterns.Mechanical approaches include confusion matrices and measurable input properties, while detailed inspection can reveal failures involving sarcasm, coordination, or subordinate clauses.
  • Disaggregated analysis: Disaggregated analysis reveals performance disparities across unitary and intersectional subgroups that aggregate metrics may conceal.
  • Counterfactual analysis: Counterfactual analysis evaluates how model outputs change after controlled changes to inputs, including changes involving sensitive identity groups or distribution shifts.
  • These analyses yield rich, detailed results that are not amenable to quick cross-system comparison, but can connect system design to aspects of the problem space.The stated purpose is to inform the next iteration of system development rather than simply identify a winner.

A.3 Ablation Testing

Ablation testing isolates the contributions of system components by removing them individually and evaluating the resulting modified systems.

  • Ablation testing removes system components one by one to isolate their contributions through evaluation of the modified system.Before deep learning, statistical NLP commonly applied ablations to feature sets to examine which information systems used.

A.4 Analysis of Model Properties

Aggregate or detailed test-item performance is only one dimension of system evaluation, particularly for judging practical feasibility.

  • Practical evaluation should consider dimensions beyond test-item performance, including energy consumption and memory and compute requirements.
Loading 2111.15366v1…