Source-linked AI summary
Toward an Evaluation Science for Generative AI Systems
Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach, Margaret Mitchell, Angelina Wang, Olawale Salaudeen, Rishi Bommasani, Deep Ganguli, Sanmi Koyejo, William Isaac
TL;DR
Generative AI is increasingly used in real-world settings, but current evaluations do not reliably anticipate performance and safety there. This paper draws on safety engineering and measurement science to propose a more mature evaluation science, emphasizing real-world applicability, iterative refinement, and institutional support. Its approach must also address generative AI’s open-endedness, non-determinism, and evolving social interactions.
Problem
Current generative-AI evaluations do not adequately anticipate or understand performance and safety in real-world deployment contexts.
Method
The paper draws lessons from safety engineering and measurement science to outline a rigorous evaluation ecosystem for generative AI systems.
Results
The paper identifies real-world applicability, iterative refinement, and institutional investment as central requirements for generative-AI evaluation.
Takeaways & Limitations
Evaluation science for generative AI should connect pre-deployment assessment, post-deployment monitoring, and shared institutional infrastructure.
Takeaways & Limitations
Generative AI’s open-endedness, non-determinism, longitudinal interactions, and costly evaluation infrastructure constrain current evaluation practices.
Abstract
from arXiv · showhide
There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluation ecosystem is insufficient: Commonly used static benchmarks face validity challenges, and ad hoc case-by-case audits rarely scale. In this piece, we advocate for maturing an evaluation science for generative AI systems. While generative AI creates unique challenges for system safety engineering and measurement science, the field can draw valuable insights from the development of safety evaluation practices in other fields, including transportation, aerospace, and pharmaceutical engineering. In particular, we present three key lessons: Evaluation metrics must be applicable to real-world performance, metrics must be iteratively refined, and evaluation institutions and norms must be established. Applying these insights, we outline a concrete path toward a more rigorous approach for evaluating generative AI systems.
1 Introduction
Generative AI is increasingly deployed in consequential real-world settings, yet its performance and safety remain poorly understood. Static benchmarks and existing evaluation practices do not adequately support real-world assurance, motivating a more mature evaluation science.
- 1 Introduction: Generative AI deployment across medicine, law, education, information technology, and social settings has exposed poorly anticipated performance and safety.Reported problems include misinformation, incorrect legal references, educational failures, and search confusion.
- 1 Introduction: Static benchmarks are not well suited to understanding the real-world performance and safety of deployed generative AI systems.Despite this mismatch, they continue to inform deployment criteria, marketing, critiques, procurement, and policy discourse.
- 1 Introduction: A mature evaluation science requires theories, testable hypotheses, reliable and generalizable measurement instruments, and iterative refinement.These properties distinguish scientific evaluation from rapidly changing, case-by-case practices.
- 1 Introduction: Leaderboards, benchmarks, and audits cannot assure system performance across different domains or user groups.The paper argues that these tools do not yet constitute the robust evaluation ecosystem needed for widespread use.
- 1 Introduction: Drawing on systems safety engineering and measurement science, the paper outlines a concrete path toward a more rigorous evaluation ecosystem.The proposed path is intended to address the gap between deployed systems and current evaluation practices.
2 Lessons from Other Fields
Other fields show that trustworthy evaluation tracks real-world risks, refines its targets and instruments over time, and depends on institutions, infrastructure, and enforceable standards.
- 2 Lessons from Other Fields: Real-world risk evaluation combines pre-deployment testing with post-deployment monitoring.Pre-deployment detection can enable earlier and cheaper mitigation, while monitoring identifies emergent harms and unexpected use.
- 2 Lessons from Other Fields: Evaluation metrics and measurement approaches must be iteratively refined as experience reveals new risks and responsibilities.Automotive safety shifted from driver-focused interventions toward seatbelts, design choices, and manufacturer responsibility.
- 2 Lessons from Other Fields: Refining measurement targets also requires refining the instruments used to capture them.Thermometer development illustrates how improved instruments can support deeper understanding, while multiple methods can provide more robust insights.
- 2 Lessons from Other Fields: Institutions can centralize testing resources and coordination for long-term, large-scale evaluation.The FDA’s development of pharmaceutical and nutrition testing regimes followed institutional investment and centralized testing efforts.
- 2 Lessons from Other Fields: Shared tools, infrastructure, and standards established through institutions can produce more thorough evaluation regimes.Transportation safety institutions combined government-mandated standards with active monitoring and reported substantial lives saved.
3 Towards an Evaluation Science for Generative AI
A generative-AI evaluation science must account for open-ended, stochastic systems and real-world social contexts. The paper therefore advocates behavioral, sociotechnical, iterative, and institutionally supported evaluation practices connected across deployment stages.
- 3.1 Unique Challenges of Generative AI: Generative AI’s open-ended use cases and non-deterministic outputs make precise measurement targets and behavior prediction difficult.Training data, model design, and user-interface choices cannot always be directly traced to downstream outputs and impacts.
- 3.1 Unique Challenges of Generative AI: Longitudinal social interactions create interaction risks that may evolve unexpectedly over time.These risks include harmful human–AI relationships and support the need for behavioral evaluation.
- 3.1 Unique Challenges of Generative AI: A behavioral approach evaluates AI system performance in different real-world settings and can connect systemic-impact evaluation with computational methods.This approach treats systems as black boxes when translating between evaluation levels.
- 3.2 Real-world applicability of metrics: Evaluation metrics should be task-specific, context-sensitive, and grounded in a broader sociotechnical lens rather than coarse claims of general intelligence.Real-world evaluation must account for factors beyond technical specifications and cannot be one size fits all.
- 3.2 Real-world applicability of metrics: Pre-deployment benchmarks should be connected to post-deployment monitoring so real-world evidence can refine and calibrate evaluation design.Naturalistic deployment interactions offer one route for improving early-stage benchmarks.
- 3.4 Establishing Institutions & Norms: Maturing evaluation requires investment in shared tools, infrastructure, institutions, and norms.These resources can support harm discovery, standards identification, evaluation guidance, and coordination across stakeholders.
- 3.4 Establishing Institutions & Norms: Open and transparent evaluation infrastructure requires substantial financial and computational investment.Running HELM once on 30 models cost USD $38,000 for commercial APIs and required 20,000 A100 hours for open models.
4 Moving Forward
Generative AI’s recent transition into widespread use has outpaced the maturity of its evaluation ecosystem. The paper argues that safety-engineering lessons and responsible prioritization can guide its development.
- 4 Moving Forward: Generative AI systems have only recently entered real-world use, so their evaluation ecosystem is not yet mature.The paper cautions that widespread deployment does not imply the elaborate safety and performance evaluations established in other fields.
- 4 Moving Forward: Evaluations cannot cover every possible use case, making prioritization and value judgments unavoidable.The paper identifies high-risk domains and deployments affecting vulnerable populations as principled priorities.
- 4 Moving Forward: Safety engineering and measurement science offer practices for anticipating failures before deployment, monitoring incidents afterward, and refining evaluations iteratively.The paper also emphasizes investing in institutions that support accessible and robust evaluation ecosystems.
- 4 Moving Forward: Generative AI’s distinctive challenges reinforce the need to create an evaluation science specific to the field.The paper presents these challenges as a reason to establish mature evaluation practices rather than an exemption from responsibility.