Source-linked AI summary
Sociotechnical Safety Evaluation of Generative AI Systems
Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, William Isaac
TL;DR
As generative AI becomes widely used and embedded, evaluating its potential harms is increasingly pressing, while capability-focused evaluations alone cannot determine safety because human and systemic factors also shape risk. The paper proposes a three-layered sociotechnical framework and reviews existing evaluations through a repository. The review identifies coverage, context, and multimodal gaps and presents practical steps and roles for addressing them.
Problem
Capability evaluations alone are insufficient to determine whether generative AI systems are safe because human and systemic factors co-determine harms.
Method
The paper proposes a three-layered sociotechnical evaluation framework and reviews current safety evaluations through a repository of existing evaluations.
Results
The review identifies three high-level gaps: insufficient coverage of several risks, rare human-interaction and systemic evaluations, and missing evaluations for multimodal AI systems.
Takeaways & Limitations
The identified gaps are tractable, with practical steps and proposed roles and responsibilities for different actors to help close them.
Takeaways & Limitations
The evaluation repository is not assumed to be comprehensive, codes modality by output modality rather than input modality, and represents a snapshot in time.
Abstract
from arXiv · showhide
Generative AI systems produce a range of risks. To ensure the safety of generative AI systems, these risks must be evaluated. In this paper, we make two main contributions toward establishing such evaluations. First, we propose a three-layered framework that takes a structured, sociotechnical approach to evaluating these risks. This framework encompasses capability evaluations, which are the main current approach to safety evaluation. It then reaches further by building on system safety principles, particularly the insight that context determines whether a given capability may cause harm. To account for relevant context, our framework adds human interaction and systemic impacts as additional layers of evaluation. Second, we survey the current state of safety evaluation of generative AI systems and create a repository of existing evaluations. Three salient evaluation gaps emerge from this analysis. We propose ways forward to closing these gaps, outlining practical steps as well as roles and responsibilities for different actors. Sociotechnical safety evaluation is a tractable approach to the robust and comprehensive safety evaluation of generative AI systems.
Reader’s guide
The paper offers reading paths tailored to different audiences, from a two-minute overview to deeper guidance for evaluators, policymakers, and researchers.
- A two-minute read focuses on the three-layered evaluation framework and figures depicting the current state of safety evaluations.It directs readers to figure 2.1 and figures 3.1–3.
- A ten-minute read adds the abstract and a skim of the framework section to the same evaluation-landscape figures.This path combines the abstract, section 2, and figures 3.1–3.
- Evaluators: Evaluators are directed toward the framework, the safety-evaluation survey, practical steps for closing gaps, and the misinformation case study.The recommended path also includes discussion of evaluation as responsible innovation and methodological limitations.
- People steering AI labs: People steering AI labs are directed to the framework, evaluation gaps, actor roles and responsibilities, methodological limitations, and the misinformation case study.This route emphasizes both organizational responsibilities and limitations of evaluation methods.
- Public policy makers: Public policy makers are directed to the framework, the current state of evaluation, and the discussion of roles and responsibilities.The recommended route prioritizes section 3 and the discussion section’s treatment of responsibilities.
- AI researchers: AI researchers are directed to the framework, practical ways forward, the case study, and the discussion of limitations and implications.The route spans sections 2, 4, the case study, and discussion.
1. Introduction
Generative AI systems are increasingly used across domains and modalities, making evaluation of their potential harms more pressing. The paper proposes a sociotechnical evaluation framework and reviews existing evaluations to identify gaps and practical responses.
- Generative AI systems are increasingly used across domains including medicine, news, politics, and social interaction.The paper describes applications spanning multiple real-world domains.
- Generative AI systems increasingly include audio, video, audiovisual, and multimodal capabilities alongside text and image generation.The paper defines multimodal systems as accepting and producing combinations of image, audio, and text.
- Generative AI systems pose risks of harm that have been mapped in taxonomies and research on individual risks or applications.The paper motivates evaluation by identifying risks across modalities and use cases.
- As generative AI becomes widely used and embedded, evaluating potential harms becomes a public safety concern and a growing priority for multiple stakeholders.The paper names AI developers, policymakers, regulators, and civil society as relevant actors.
- Safety evaluation measures AI-system performance or impact through exploratory and directed approaches, including qualitative investigations of actual use.Evaluation also involves technical and normative decisions about what to measure and what counts as good performance.
- Current safety evaluations are heterogeneous and ad hoc, making them difficult to compare or reproduce and potentially missing important risks.The paper argues for a more systematic and standardised approach to meaningful, comparable, and comprehensive evaluation.
- The paper contributes a three-layered sociotechnical framework and a review with a repository of existing evaluations, gap analysis, and proposed stakeholder responsibilities.It also presents practical steps toward closing the identified gaps.
2. Framework for sociotechnical AI safety evaluation
The paper proposes a three-layered sociotechnical framework that evaluates capabilities alongside human interaction and systemic impacts, because technical evaluation alone cannot determine safety. The layers progressively add context and can be evaluated simultaneously to support comprehensive safety evaluation.
- Technical evaluation alone is insufficient because human and systemic factors co-determine risks of harm.
- The framework structures evaluation around capability, human interaction, and systemic impact layers.These layers target different aspects of an AI system and progressively add context needed to assess whether capabilities relate to actual harm.
- Layer interactions: The framework treats layer boundaries as gradual and interactive, allowing methods such as adversarial testing to assess capabilities and user friction simultaneously.Observations at one layer may foreshadow effects at another, such as disparate performance across user groups indicating potential systemic impacts.
- Capability evaluation: Capability evaluation tests technical components, system behaviour, outputs, training data, and processes for indicators of potential downstream harm.Examples include harmful stereotypes, factual errors, performance differences across languages or groups, efficiency metrics, training-data representativeness, and learned associations.
- Human interaction evaluation: Capability evaluation remains critical but requires contextual assessment of who uses a system, for what purpose, and under which circumstances.The human interaction and systemic impact layers provide this additional context.
3. Current state of sociotechnical safety evaluation
The survey maps existing generative-AI safety evaluations across harm areas, modalities, and evaluation layers, revealing coverage, context, and multimodal gaps. Evaluations are concentrated at the capability and text-output levels, while broader sociotechnical and multimodal coverage remains limited.
- Survey approach: The survey maps identified evaluations by risk area, AI-system modality, and the three evaluation layers to characterize the current safety-evaluation landscape.The mapping is presented as a snapshot and is used to identify evaluation gaps.
- Harm taxonomy: The taxonomy organizes generative-AI harms into six areas: representation and toxicity, misinformation, information and safety, malicious use, autonomy and integrity, and socioeconomic and environmental harms.It synthesizes prior literature into a holistic taxonomy for mapping safety evaluations.
- Evaluation gaps: Three gaps emerge: low coverage of several risks, rare human-interaction and systemic evaluations, and missing evaluations for multimodal AI systems.Existing evaluations cluster at the capability layer and primarily target text output.
- Coverage gap: Coverage is especially scarce for information and safety harms, autonomy and integrity harms, and socioeconomic and environmental harms.The authors caution that available evaluation counts are insufficient to establish comprehensive coverage.
- Coverage gap: Even representation-harm evaluations cover narrow subspaces: 17% assess binary gender and occupation bias, while 60 evaluations target text modalities only.Other potential axes include ability status, age, religion, nationality, and social class.
- Context gap: Capability evaluations do not account for contextual factors that co-determine harm, so human-interaction and systemic-impact analyses are needed alongside capability testing.The framework therefore adds layers that progressively incorporate relevant context.
- Multimodal gap: Most evaluations assess text output; only four publicly documented evaluations target audio, and none were found for video.Few evaluations address image outputs or combinations of text and image.
- Multimodal gap: Multimodal outputs can create novel or compound risks, requiring evaluations of individual modalities and their compositions rather than text-only assessment.The paper notes that harm manifestations can differ across modalities and combinations.
4. Closing evaluation gaps
The paper identifies substantial sociotechnical safety-evaluation gaps and proposes practical approaches for constructing, extending, and validating evaluations across three layers. It emphasizes that operationalising complex harms into measurable constructs requires careful attention to validity, while multimodal and model-driven methods introduce additional limitations.
- Closing evaluation gaps: The paper proposes closing evaluation gaps by constructing novel evaluations, extending existing evaluations to generative AI systems, and clarifying roles and responsibilities.The proposed work includes a general construction pipeline, operationalisation of harms, and practical avenues for extending existing evaluations.
- Evaluation as a process: Evaluation proceeds by selecting a target, operationalising it into a metric, obtaining a measurement, and judging the outcome.Each step includes both technical and normative elements.
- Operationalising risks: Risks of harm require operational definitions because they are latent constructs that cannot be directly observed through a single test or metric.Operationalisation maps observable metrics or concepts to latent constructs so measurements can provide insight into the target risk.
- Three-layer evaluation: Comprehensive assessment requires complementary evaluation at the capability, human interaction, and system layers, with harm constructs operationally defined at each layer.Different aspects of a risk may require different metrics at different layers.
- Validity: Narrow operational definitions can produce internal and external validity failures when metrics capture only part of a complex harm or fail to generalize beyond the tested setting.The paper illustrates this problem with a stereotyping measure based on word-pair associations that included innocuous pairs such as “Norwegian” and “salmon”.
- Validity: Recommended validity practices include grounding operationalisations, incorporating diverse perspectives, cross-validating related evaluations, and preserving interpretability when aggregating results.Aggregating multiple tests can capture several facets of harm but may obscure item-level validity failures.
5. Discussion
Technical and behavioral evaluations alone cannot determine generative AI safety because harms emerge through interactions with human and systemic context. The paper therefore advocates a three-layered sociotechnical approach, shared responsibilities, parallel evaluation, and complementary governance while recognizing evaluation’s limits.
- Benefits of a sociotechnical approach: Technical components and system behavior are important but insufficient for determining safety because potential harms are felt outside technical evaluations.Capabilities can predict harm risk but remain a proxy for actual downstream impacts.
- Benefits of a sociotechnical approach: The three-layered framework adds capability, human interaction, and systemic impact evaluations to account for context that determines system safety.The layers are not contingent or sequential; they can be conducted simultaneously and asynchronously.
- Roles and responsibilities: Responsibility for comprehensive evaluation is shared, with developers focused on system capabilities, application developers on human interaction, and third parties on systemic impacts.Actors are best placed according to expertise, autonomy, infrastructure, and domain knowledge.
- Roles and responsibilities: Proprietary unreleased systems may require novel infrastructure, incentives, standardized approaches, and reliable safety assurances for capability and human interaction testing.These requirements can make evaluations difficult for evaluators and developers to coordinate.
- Limits of evaluation: Evaluation cannot identify every harm because it covers only selected manifestations, may miss small or poorly understood effects, and struggles with distributed impacts and complex constructs.The authors recommend governance for remaining uncertainties, post-deployment incident monitoring, and swift recourse mechanisms.
- Evaluation gaps: Current evaluations have a pressing collective toolkit gap, are often ad hoc and late, and require concerted action to address risks from increasingly capable multimodal models.The paper also argues that mixed-methods practices are needed because high-frequency evaluations have validity constraints for complex real-world harms.
6. Conclusion
The paper proposes a sociotechnical approach that expands safety evaluation beyond capabilities to human and systemic context. It surveys existing evaluations, identifies major gaps, and offers a pragmatic roadmap for addressing them.
- The framework expands evaluation from capability testing to the context of system use and broader impacts on embedded structures.
- The survey identifies gaps in specific risk areas, non-text modalities, and evaluations addressing human and broader systemic context.
- The paper provides a pragmatic roadmap combining available methods, tactical extensions, and visions for evaluating misinformation, representation risks, and dangerous information.
A.1. Taxonomy of harm
The taxonomy organizes generative AI harms across representation, information and safety, human autonomy and integrity, and socioeconomic and environmental domains. It pairs each risk area with definitions and examples, distinguishing observed from anticipated harms.
- Representation & Toxicity Harms: Representation and toxicity harms include unfairly representing identities or perspectives, toxic content, and harmful stereotyping.Examples include gendered occupational associations and generated gruesome, abusive, or hateful imagery.
- Misinformation and Information & Safety Harms: Misinformation and information-safety harms include false beliefs, polluted information ecosystems, privacy infringement, and dissemination of dangerous information.Examples include synthetic videos prompting panic, leaked payment information, and instructions for creating novel biohazards.
- Malicious Use: Malicious-use harms include influence operations, fraud, defamation, and security threats enabled by generated or manipulated media and code.Examples include election-oriented false news, voice impersonation scams, false attributions, and code for hacking government systems.
- Human Autonomy & Integrity Harms: Human autonomy and integrity harms include non-consensual identity use, persuasion and manipulation, overreliance, skill atrophy, and exploitative appropriation.Examples include non-consensual deepfakes, assistants persuading self-harm, and reduced critical thinking from excessive use.
- Socioeconomic & Environmental Harms: Socioeconomic and environmental harms include unequal access and performance, precarity, environmental impacts, creative homogenization, and exploitative data or labor practices.Examples span unequal hiring pathways, increased carbon emissions, lower creative-professional pay, and toxic-content exposure for annotators.
- The taxonomy distinguishes real-world examples marked with (*) from hypothetical anticipated risks marked with (±).
A.2.1. Capabilities layer
The capabilities layer commonly uses benchmarks and human annotation to measure model outputs and performance. These methods support repeatable, frequent comparisons but have validity, representativeness, cost, and interpretability limits.
- Human annotation: Human annotation supplies judgments for outputs whose offensiveness, misleadingness, or social harm depends on human assessment.It is commonly treated as ground truth, although some harms also have algorithmic measures.
- Human annotation: Human annotation is costly and time-intensive, can expose annotators to harm or precarious work, and may be difficult to source when specialized expertise is needed.
- Human annotation: Annotated datasets can be biased by unrepresentative pools, quality-reducing incentives, interface design, rater demographics, and aggregation that suppresses minority disagreement.
- Benchmarking: Capability benchmarks map model outputs to predefined tasks, measuring performance against clear intended outcomes rather than exploratory output patterns.
- Benchmarking: Benchmarks support test-retest reliability, frequent model-progress tracking, and cross-model comparisons when metrics are shared.They can be cost-effective and time-effective for large AI-developing organizations despite computational and financial costs.
- Benchmarking: Benchmark results can guide AI design and responsible decisions, but held-out use matters because optimization can trigger Goodhart’s Law.
- Limitations: Psychology-inspired tests may lack validity for AI because assumptions about human or animal minds, life cycles, memory, learning, and embodiment may not hold.
- Limitations: Benchmarks may be too small, narrow, or biased, cannot capture complex social risks, and have limited external validity for novel instances.Aggregating tests into one score can also hide disparate performance across groups; dynamic benchmarks are proposed as one response.
A.2.2. Human interaction layer
The human interaction layer evaluates how people use and experience AI systems, combining controlled studies, real-world observation, and mixed methods. These evaluations address interaction-related harms but face limits in causal inference, generalisability, scale, and research ethics.
- A.2.2. Human interaction layer: Controlled studies isolate variables and investigate outcomes or mechanisms through psychology, human–computer interaction, or behavioural economics methods.They can examine potential impacts of interactions with general-purpose technologies.
- A.2.2. Human interaction layer: User research evaluates user needs, behaviours, functionality, and externalities in controlled or real-world settings.Methods include behavioural experiments, interviews, talk-out-loud studies, and surveys.
- A.2.2. Human interaction layer: User studies may lack the ethics scrutiny applied to psychology experiments, and some have adversely affected participant well-being.This creates an additional safety concern in the conduct of human interaction evaluations.
- A.2.2. Human interaction layer: User testing is increasingly necessary because human interaction can involve overtrust, overreliance, anthropomorphism, emotional harm, and disparate functionality across user groups.Historically, user testing focused more on product development than safety evaluation, and its results were often not publicly disclosed.
- A.2.2. Human interaction layer: Passive monitoring can reveal interaction patterns and effects during or beyond human–AI interactions, including risks across a user journey.It can be combined with interviews, surveys, active interventions, quasi-experiments, and longitudinal analysis.
- A.2.2. Human interaction layer: Passive monitoring rarely isolates causal effects, while experiments trade real-world realism and scale for stronger insight into causal mechanisms.Small studies may miss harms with small effect sizes, and user testing findings often do not extend beyond the specific application and context studied.
A.2.3. Systemic impact layer
The systemic impact layer evaluates effects of AI systems on broader economic, environmental, and social structures using pilots, impact assessments, observations, forecasts, and simulations. These approaches can examine real-world and downstream effects, but causal evidence and generalisation remain constrained.
- Staged release and pilot studies: Staged releases and pilot studies deploy AI systems in controlled real-world settings to monitor systemic effects such as care provision and productivity.They can also support experiments comparing outcomes across organisations with different levels of AI deployment.
- Staged release and pilot studies: Pilot studies may not generalise to novel contexts, and small-scale results can be overturned by equilibrium effects during large-scale adoption.These weaknesses limit how confidently local or small experiments predict wider systemic outcomes.
- Impact assessments: Impact assessments can be prospective or retrospective and track potential effects on broader economic, environmental, or social structures.They may use guiding questions, system-level indicators, ethnographic methods, expert views, case studies, and incident aggregation.
- Impact assessments: Interviewing different groups can reveal uneven systemic impacts, including increased negative emotions for low-skilled employees but increased creativity for high-skilled employees.This example illustrates distributional differences across groups.
- Impact assessments: Impact assessments rely on qualitative or observational data about complex, prolonged processes, making causal evidence difficult to generate.Natural experiments can partially address this difficulty but restrict study scope and require strong assumptions about technology adoption.
- Forecasts and simulations: Forecasts and simulations map anticipated capabilities against job tasks or use comparative technologies to estimate downstream labour-market impacts.Loose analogies may serve as heuristics when highly similar technologies do not exist, but should be used cautiously.
- Forecasts and simulations: Forecast results are uncertain because ambiguous choices about defining and quantifying model capabilities affect estimates of economic exposure to automation.Transparent data sources and methods are therefore critical to interpreting such evaluations.
A.3. Case study: Misinformation
Misinformation illustrates why safety evaluation must connect AI capabilities with human interpretation and broader social context. The case study distinguishes false content from intentional deception and shows that harm depends on belief, stakes, timing, and societal consequences.
- A.3. Case study: Misinformation: Misinformation is the often-unintentional spread of false, inaccurate, or misleading information, distinct from intentionally deceptive disinformation.Generative AI can produce realistic but factually incorrect or misleading text, images, audio, and video.
- A.3. Case study: Misinformation: Factually inaccurate outputs may be harmless in creative contexts but harmful when they prompt action on false beliefs in legal, medical, or other high-stakes settings.People may also struggle to distinguish synthetic from human-generated content.
- A.3. Case study: Misinformation: Synthetic misinformation may erode public trust, increase uncertainty about what to believe, and undermine shared public knowledge and evidence authentication.These concerns extend beyond individual interactions to broader societal repercussions.
- A.3. Case study: Misinformation: Comprehensive misinformation evaluation requires measuring concepts across all three sociotechnical layers, linking latent harms to intermediate concepts and concrete measures.Factuality is presented as one concept narrower than misinformation and broader than a concrete metric such as FID scores.
A.3.1. Capability
Capability evaluations assess whether AI systems generate factually accurate, verifiable, realistic, and contextually relevant outputs. Existing metrics and benchmarks provide partial evidence, while human expertise and stronger multimodal datasets remain necessary for evaluating misinformation risk.
- A.3.1. Capability: Capability-layer evaluations assess the factual accuracy of AI system outputs through benchmarks and other tests, with most existing work focused on text rather than image, audio, or video.Textual data availability contributes to this modality imbalance.
- A.3.1. Capability: Automatic fact verification, source attribution, and factual-knowledge benchmarks test whether outputs are supported by reliable sources or contain false statements.Examples include FActScore, RARR, Wiki-FACTOR, News-FACTOR, FactualityPrompt, and TruthfulQA.
- A.3.1. Capability: Large, up-to-date multimodal verification datasets remain a major roadblock, and general alignment metrics are unsuitable for detecting AI-generated audiovisual content.Existing datasets support deepfake detection, but coverage remains limited.
- A.3.1. Capability: Human factuality evaluation is laborious, depends on raters and verifiable evidence, and can become outdated as contested or emerging knowledge changes.Benchmarks capture a limited operationalisation of truthfulness, while source-attribution approaches may ignore source quality and trustworthiness.
- A.3.1. Capability: Factuality is insufficient as a proxy for misinformation because harmless implausibility may be ignored, whereas believable false content can deceive audiences.Credibility depends on realism, prior beliefs, knowledge, source trust, and viewing context.
- A.3.1. Capability: Realism and perceptual-quality metrics, including FID and Inception scores, compare generated outputs with real references, but fidelity scores may correlate poorly with human quality judgments.Human evaluations commonly measure perceived realism and perceptual quality, while some benchmarks test whether people can distinguish real from AI-generated images.
- A.3.1. Capability: Misinformation risk also depends on societal relevance, such as political salience, election timing, content domain, and the likelihood that a specific output will be believed.Experts may help establish criteria for evaluating high-risk multimodal outputs.
A.3.2. Human interaction
Human-interaction evaluations examine whether generative AI outputs deceive or persuade people, including effects on beliefs, attitudes, and behaviour. Systemic impacts require deployment-scale measurement, while detection and watermarking methods remain vulnerable to limited generalisation and manipulation.
- Deception: Human-interaction evaluations test whether users identify synthetic or misleading outputs and whether AI-generated content deceives them.Studies examine synthetic-versus-human discrimination, misinformation identification, and cues such as facial expressions and gestures.
- Persuasion: AI-generated outputs can influence users’ beliefs, attitudes, and behaviour, including political views and policy attitudes.Evaluations use interaction or exposure experiments to assess persuasion and the development of false beliefs about external reality.
- Persuasion: Human-centred experiments can examine cognitive, social, and demographic factors shaping receptivity to misinformation and false beliefs.Reported factors include partisanship, cognitive biases, trust in information sources, and repeated exposure.
- Systemic impacts: Societal impacts such as erosion of public trust and information pollution can be empirically measured only after deployment, once adoption and time allow effects to manifest.Proposed measures include population-level shifts in trust and the prevalence of false or misleading synthetic content in the public domain.
- Systemic impacts: Detection tools and watermarks often struggle across languages and transformed media, and can be defeated by resizing, cropping, or format changes.These techniques therefore require continuous updating to prevent misuse.