Source-linked AI summary
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
Hanna Wallach, Meera Desai, A. Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Alexandra Olteanu, Nicholas Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vaughan, Matthew Vogel, Hannah Washington, Abigail Z. Jacobs
TL;DR
GenAI evaluations measure abstract, contested concepts, making it difficult to determine what instruments measure and whether their results are valid. The paper adapts social-science measurement theory into a four-level framework that separates conceptual and operational decisions, broadens stakeholder participation, and organizes validity interrogation. The framework standardizes measurement without guaranteeing that evaluations will meet every intended goal.
Problem
GenAI evaluation concepts are abstract and contested, making it difficult to know precisely what measurement instruments measure, why, and whether resulting measurements are valid.
Method
The paper proposes a four-level framework grounded in social-science measurement theory, linking concepts, instruments, measurements, and validity interrogation.
Results
The framework separates conceptual from operational debates and enables stakeholders with different perspectives to participate in conceptual measurement decisions.
Takeaways & Limitations
The framework provides a cohesive way to standardize measurement and clarify precisely what GenAI evaluation instruments are and are not measuring.
Takeaways & Limitations
Adopting the framework is challenging for researchers and practitioners with limited time, budgets, or other resources, and it is not a panacea.
Abstract
from arXiv · showhide
The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from learning from and drawing on the social sciences when developing and using measurement instruments for evaluating GenAI systems. Specifically, our position is that evaluating GenAI systems is a social science measurement challenge. We present a four-level framework, grounded in measurement theory from the social sciences, for measuring concepts related to the capabilities, behaviors, and impacts of GenAI systems. This framework has two important implications: First, it can broaden the expertise involved in evaluating GenAI systems by enabling stakeholders with different perspectives to participate in conceptual debates. Second, it brings rigor to both conceptual and operational debates by offering a set of lenses for interrogating validity.
1. Evaluating GenAI Systems
Evaluating GenAI systems requires measuring capabilities, behaviors, and impacts that are often abstract and contested. The paper argues that ML should draw on social-science measurement theory to make evaluations more valid and rigorous.
- Motivation: GenAI evaluations inform decisions about system use, deployment, and redesign by measuring capabilities, behaviors, and impacts.These measurements can concern mathematical reasoning, training-data regurgitation, or users feeling harmed.
- Motivation: GenAI measurement is harder than traditional ML measurement because systems accept varied inputs, produce diverse outputs, support many use cases, and affect society.The relevant concepts are often abstract, indirectly measured, and contested across use cases, cultures, and languages.
- Problem: Existing instruments make it difficult to know precisely what they measure, why they measure it, and whether their measurements are accurate or useful.The paper frames this as a validity problem arising from abstract concepts and common ML measurement practices.
- Position: The paper positions GenAI evaluation as a social-science measurement challenge and urges ML researchers to learn from social-science measurement practice.Social scientists have long measured abstract, contested concepts such as ideology, democracy, media bias, and framing.
- Position: The proposed framework standardizes the measurement process without transferring human-designed social-science instruments to GenAI systems.It is presented as applicable to both GenAI and traditional ML systems while preserving existing instruments and exposing gaps in current practice.
2. Related Work
Related work identifies serious limitations in GenAI measurement instruments and develops partial responses through psychometrics, responsible AI, and benchmark consolidation. The paper unifies these strands with a broader framework for standardizing measurement.
- Critiques of measurement instruments: Current GenAI measurement instruments exhibit conceptual confusion, insufficient validity interrogation, and conflicting measurements across benchmarking, red teaming, and real-world evaluations.These critiques motivate a broader account of measurement beyond benchmark construction.
- Parallels to psychometric tests: Psychometric work highlights structural similarities between GenAI benchmarks and psychological tests while emphasizing validity interrogation.The similarities include items, scoring rubrics, and aggregation functions.
- Responsible AI: Responsible AI research similarly emphasizes that abstract, contested measurements reflect different meanings and understandings and require validity interrogation.These ideas have mostly been used to critique instruments rather than standardize measurement as a whole.
- Benchmark consolidation: Benchmark consolidation improves model comparability and evaluation reproducibility by standardizing conditions such as datasets and model parameters.These efforts primarily address benchmarking rather than other measurement approaches.
- Synthesis: The paper argues that a cohesive framework can unify related work and help practitioners avoid measurement limitations by design.The proposed path forward responds to slow or unsuccessful uptake across ML subcommunities.
3. A Measurement Framework for GenAI
The framework adapts social-science measurement theory into four linked levels and processes for evaluating GenAI concepts. It separates conceptual decisions from operational decisions, broadens stakeholder participation, and treats validity as context-dependent evidence.
- Four-level framework: The framework distinguishes background concepts, systematized concepts, measurement instruments, and measurements, linked by systematization, operationalization, application, and interrogation.Systematization narrows a broad concept, operationalization develops instruments, application obtains measurements, and interrogation evaluates validity.
- Applications: The framework applies across diverse tasks, including stereotype prevalence, mathematical reasoning, memorization, and harmful-prompt refusal.These examples span capabilities, behaviors, and impacts in deployed or modeled GenAI systems.
- Separating systematization and operationalization: Separating systematization from operationalization distinguishes debates about what is measured from debates about how it is measured.This separation makes conceptual meanings and operational validity independently interrogable.
- Separating systematization and operationalization: The separation enables stakeholders with different perspectives to participate in conceptual debates about which meanings and understandings measurements should reflect.Relevant stakeholders include developers, policymakers, customers, users, and marginalized communities.
- Emphasizing interrogation: Validity must be interrogated as evidence supporting particular interpretations and uses within a measurement context, because validity may not transfer across contexts.The framework recommends lenses including face, content, convergent, discriminant, predictive, hypothesis, and consequential validity.
- Emphasizing interrogation: Determining what counts as sufficiently valid is itself difficult when abstract concepts lack directly observable and universally agreed-upon labels or scores.This difficulty is a central boundary on validity assessment rather than a problem resolved automatically by the framework.
4. Using the Measurement Framework
The framework separates conceptual systematization from operationalization and uses continual validity interrogation to connect abstract concepts with observable phenomena and measurements. A running example shows how definitions, indicators, aggregation rules, and instruments are developed and tested.
- Systematization: Systematization narrows a broad background concept into an explicit definition that specifies precisely what will be measured and why.
- Systematization: Stakeholders with different perspectives can participate in conceptual debates through the systematization process.
- Systematization: In the stereotyping example, four linguistic patterns define the concept, and any one pattern is sufficient for a text to count as stereotyping.
- Interrogation: Face, content, and consequential validity support continual interrogation of the systematized concept and can reveal contested definitions or consequences of measurement.
- Operationalization: Operationalization defines indicators, specifies their aggregation, and develops measurement instruments that obtain the resulting measurements.
- Interrogation: Validity should be rigorously interrogated because high-level annotation guidelines and few-shot examples do not clearly establish what ML measurement instruments measure or whether measurements are valid.
5. Adopting the Measurement Framework
The paper recommends adopting the framework through conceptual and operational debate, transparent sharing of concepts and validity evidence, and resource-sensitive implementation. It presents partial adoption as worthwhile while acknowledging practical and organizational barriers.
- Engage in conceptual debates: Researchers should involve stakeholders with different perspectives, draw on relevant disciplines, and use validity lenses during conceptual debates.
- Engage in operational debates: Researchers should interrogate the validity of instruments and measurements, distinguish background from systematized concepts, and re-interrogate validity in new contexts.
- Share the systematized concept: Sharing the systematized concept alongside instruments and measurements clarifies what is being measured and helps identify apples-to-oranges comparisons.
- Share evidence of validity: Sharing evidence both for and against validity helps others make informed decisions about using measurement instruments or their results.
- Resource-sensitive adoption: The framework is an ideal whose full adoption may challenge researchers with limited time, budgets, or other resources, but partial adoption can still contribute to changing evaluation practice.
- Resource-sensitive adoption: Separating systematization from operationalization and interrogating validity are the most important actions, with comprehensiveness adjusted to available resources and purpose.
- Organizational support: Changing GenAI evaluation practice will be challenging, so organizations should provide resources and incentives and remove organizational barriers to adoption.
6. Alternative Views
The paper addresses alternatives to its position by arguing that GenAI evaluations require a structured measurement framework because their concepts are abstract, socially intertwined, and increasingly consequential. It proposes drawing on social-science measurement without transferring human measurement instruments or anthropomorphizing systems.
- Existing evaluations have serious limitations, and a standardized process can clarify when and why measurements are comparable.
- The framework offers a structured way to state and interrogate assumptions, unifying and reinterpreting existing ML practices.
- GenAI concepts are abstract and deeply intertwined with people and society, making social-science measurement frameworks relevant to evaluation.
- The framework does not imply that GenAI systems should be anthropomorphized or evaluated with instruments designed for humans.
- Separating systematization from operationalization parallels separations that supported advances in computer architecture, internet measurement, and programming languages.
7. Conclusion
The paper concludes that GenAI evaluation should mature into a rigorous science of measurement. It advocates a social-science-based framework to standardize measurement and clarify debates about evaluating capabilities, behaviors, and impacts.
- The paper positions GenAI evaluation as a social science measurement challenge and advocates learning from social sciences when developing measurement instruments.
- Its four-level framework is grounded in social-science measurement theory and addresses concepts related to GenAI capabilities, behaviors, and impacts.
- The framework is intended to clarify existing ML measurement debates and promote more rigorous measurement practices.
Impact Statement
The paper presents the framework as a way to improve conceptual clarity, expertise, and measurement validity, while emphasizing that it is not sufficient by itself to improve downstream practice or policy. It also supports qualitative and quantitative measurement and may reveal shortcomings in evaluations.
- The framework does not inevitably improve how GenAI systems are developed, deployed, used, or regulated.
- Improving evaluations must be accompanied by sustained efforts to inject research meaningfully into policymaking and practice.
- The framework supports qualitative and quantitative measurement instruments, although measurements themselves are quantitative.
- The framework is not a panacea: it standardizes measurement and clarifies what instruments do and do not measure, but may reveal remaining shortcomings.
A. Terminology
The terminology distinguishes the object, concept, context, instruments, approach, process, and evaluation, then separates systematization from operationalization. Validity is interrogated through multiple lenses across conceptual and operational debates.
- Core entities: The object of interest is the GenAI system being evaluated, while the concept of interest concerns its capabilities, behaviors, or impacts.
- Core entities: The context of interest specifies where the concept is measured, including adversarial use or real-world deployment contexts.
- Core entities: Measurements are quantities on nominal, ordinal, interval, or ratio scales that reflect a concept exhibited by an object in a context.
- Measurement architecture: Measurement instruments are operational procedures and artifacts, including datasets, classifiers, annotation guidelines, scoring rubrics, and aggregation functions.
- Measurement architecture: A measurement approach combines a high-level strategy with a scope defining what will and will not be measured.
- Measurement architecture: The measurement process obtains measurements systematically, whereas the broader evaluation process uses information to make and justify claims about an object.
- Concept development: Indicators represent observable phenomena connected to the concept, and their values are obtained and aggregated through measurement instruments.
- Concept development: Systematization narrows a contested background concept into an explicit systematized concept, while operationalization develops instruments from that formulation.
B. Lenses of Validity
The paper presents validity as multiple sources of evidence for evaluating both conceptual definitions and measurement instruments. These lenses support scrutiny of what is measured, how it is measured, and what consequences measurement produces.
- Face validity: Face validity asks whether a concept, instrument, or resulting measurement looks reasonable, but it is subjective and requires less subjective evidence.Anyone can interrogate face validity, including concept systematizers and instrument developers.
- Content validity: Content validity assesses whether a systematized concept captures salient aspects of its background concept and whether instruments align with the concept’s substance and structure.Its substantive and structural facets address observable phenomena and their relationships to the concept.
- Convergent validity: Convergent validity compares measurements with similar measurements from other already validated instruments and can inform both conceptual and operational debates.Different systematized concepts can make dissimilar results difficult to attribute to conceptual definition, operationalization, or both.
- Additional validity lenses: Discriminant, hypothesis, and predictive validity provide evidence by comparing dissimilar concepts, testing established hypotheses, or predicting related external phenomena.Failure in hypothesis or predictive validity can indicate systematization issues, operationalization issues, or both.
- Consequential validity: Consequential validity examines intended and unintended consequences across measurement processes, concepts, instruments, and measurements, including societal, ethical, and cultural effects.The paper describes it as the widest-ranging validity lens and connects it to the possibility that optimization can diminish validity over time.
- Applicability: The framework applies to diverse measurement tasks, including stereotyping, mathematical reasoning, memorization, and harmful-prompt refusal, while the examples are illustrative rather than comprehensive.The authors use four very different tasks to demonstrate the framework’s applicability.
D. Interrogation: Operationalization (Cont.)
The appendix applies validity interrogation to measuring stereotypical text in chatbot outputs and uses the example to distinguish discriminant and hypothesis validity. It also situates the framework within broader debates about GenAI measurement.
- Discriminant validity: Discriminant validity for stereotype measurement compares a judge LLM’s results with validated measurements of dissimilar concepts such as hostile text or negative sentiment.These comparison concepts are related to, but not identical with, stereotypical text.
- Hypothesis validity: Hypothesis validity can be tested by checking whether inputs independently designed to elicit stereotypes produce chatbot outputs with higher measured stereotype prevalence.The eliciting inputs are assumed to have been confirmed through other methods.
- Case-study context: The framework is illustrated through a case study on measuring whether GenAI models encode exact or near-exact copies of training data, a concept called memorization.The case study addresses ongoing debates about memorization.
E.1. Systematization
The memorization case study shows that systematization requires explicit decisions about what counts as memorization, while commonly measured extraction and regurgitation are distinct concepts. These choices materially shape the resulting measurements.
- Conceptual distinctions: Memorization, extraction, and regurgitation are distinct concepts that are often conflated, creating conceptual confusion in GenAI evaluation.Memorization concerns what is encoded in model parameters; extraction concerns successful prompting, while regurgitation concerns outputs containing training-data copies.
- Systematization: Systematizing memorization requires decisions about what constitutes a training-data piece and how to define exact or near-exact copying across modalities.The paper argues that foregrounding these decisions would bring greater clarity to ongoing debates.
- Implications: The case demonstrates that explicitly interrogating systematization and operationalization decisions can clarify what memorization measurements mean and how they should be interpreted.The paper presents these decisions as central to understanding the resulting measurements.
- Measurement implications: Measurements of regurgitation or extraction likely underestimate memorization because models may memorize training-data pieces that they do not generate as exact or near-exact outputs.Generating a copied training-data piece is strong evidence of memorization but does not capture all memorized pieces.
- Operationalization: Instrument development requires choices about training-data proxies, piece length, and prompting methodology, each of which can significantly influence measurements.Discoverable extraction samples 2k-token pieces, splits them into k-token prefixes and suffixes, and compares suffixes with outputs generated from prefixes.
E.2.1. INTERROGATION: OPERATIONALIZATION
Validity lenses reinterpret prior findings on memorization measurement by testing whether extraction reflects encoded training data rather than happenstance and whether it aligns with expected model behavior.
- Face and discriminant validity: 50 tokens is accepted as the norm for extraction studies because 10 tokens is too short to confidently rule out happenstance generation.The comparison concerns the number of training-data tokens used when assessing whether outputs reflect encoded data.
- Discriminant validity: Probabilistic discoverable extraction studies discriminant validity by comparing extraction from training data with exact or near-exact generation from unseen test data.Unseen test data cannot have been memorized and therefore cannot be extracted, making it a comparison for happenstance generation.
- Convergent validity: Recent work finds that extraction measurements correlate with corresponding suffix perplexities, supporting convergent validity for probabilistic discoverable extraction.The correlation matches the expected relationship between extraction and suffix perplexity.