Source-linked AI summary
Ethical and social risks of harm from Language Models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, Iason Gabriel
TL;DR
Large-scale language models create a broad ethical and social risk landscape that requires structured understanding for responsible innovation. This report synthesizes multidisciplinary expertise into a taxonomy of 21 risks across six areas, examines their origins and mitigations, and emphasizes organizational responsibility and collaboration. It presents the taxonomy as a starting point for responsible decision-making and mitigation, while leaving benefits analysis and some evolving risks outside its scope.
Problem
The report addresses the need for an in-depth, structured understanding of ethical and social risks associated with large-scale language models to support responsible innovation.
Method
The authors synthesize varied disciplinary expertise and sources to construct a taxonomy of language-model risks, analyze their origins and manifestations, and discuss mitigation approaches.
Results
The report identifies 21 risks organized into six risk areas and aims to make them actionable for organizational decisions, public discussion, and mitigation work.
Takeaways & Limitations
Responsible mitigation requires collaboration, inclusive participation, and a broad view of risks so that addressing one harm does not aggravate another.
Takeaways & Limitations
The report focuses on anticipating and structuring risks rather than analyzing benefits or performing a full ethical cost-benefit evaluation, and its taxonomy is a time-bound starting point.
Abstract
from arXiv · showhide
This paper aims to help structure the risk landscape associated with large-scale Language Models (LMs). In order to foster advances in responsible innovation, an in-depth understanding of the potential risks posed by these models is needed. A wide range of established and anticipated risks are analysed in detail, drawing on multidisciplinary expertise and literature from computer science, linguistics, and social sciences. We outline six specific risk areas: I. Discrimination, Exclusion and Toxicity, II. Information Hazards, III. Misinformation Harms, V. Malicious Uses, V. Human-Computer Interaction Harms, VI. Automation, Access, and Environmental Harms. The first area concerns the perpetuation of stereotypes, unfair discrimination, exclusionary norms, toxic language, and lower performance by social group for LMs. The second focuses on risks from private data leaks or LMs correctly inferring sensitive information. The third addresses risks arising from poor, false or misleading information including in sensitive domains, and knock-on risks such as the erosion of trust in shared information. The fourth considers risks from actors who try to use LMs to cause harm. The fifth focuses on risks specific to LLMs used to underpin conversational agents that interact with human users, including unsafe use, manipulation or deception. The sixth discusses the risk of environmental harm, job automation, and other challenges that may have a disparate effect on different social groups or communities. In total, we review 21 risks in-depth. We discuss the points of origin of different risks and point to potential mitigation approaches. Lastly, we discuss organisational responsibilities in implementing mitigations, and the role of collaboration and participation. We highlight directions for further research, particularly on expanding the toolkit for assessing and evaluating the outlined risks in LMs.
Reader’s guide
The report is organized into three segments: an introduction, a taxonomy of language-model harms, and a discussion of causes, mitigations, and future research. Sections can be read independently or together, with reading paths tailored to different time commitments and audiences.
- The report has three segments: Introduction, Classification of harms from language models, and Discussion and Directions for future research.
- Individual sections can be read independently or together.
- A one-minute route is to study Table 1 for a high-level overview of the risks considered.
- A ten-minute route combines the Abstract and Table 1 with skimming bold text in the harms classification and future-research sections.
- Readers actively working on language models are encouraged to focus on bold text and risks directly related to their work and interests.
1. Introduction
The report proposes a multidisciplinary taxonomy to structure ethical and social risks associated with language models and support responsible decision-making and mitigation. It focuses on anticipating and organizing risks rather than evaluating benefits, while recognizing that the taxonomy is time-bound and excludes several domains.
- Purpose and contribution: The report aims to underpin organizational decisions, inform public discussion, and guide mitigation work by making language-model risks actionable.
- Risk taxonomy: The authors organize 21 risks into six areas spanning discrimination, information hazards, misinformation, malicious uses, human-computer interaction, and automation, access, and environmental harms.
- Risk analysis: Each risk is discussed through its harm, empirical examples, additional considerations, and a fictitious example of how it may manifest.
- Risk analysis: The taxonomy covers language models generally, marks risks as anticipated or observed, and supports foresight about their development and use.
- Approach: The report draws on varied disciplinary expertise and sources, while emphasizing inclusive dialogue with affected communities and the wider public.
- Scope and limitations: The report does not provide a full ethical evaluation or cost-benefit analysis, and it does not survey beneficial applications or use cases comprehensively.
- Scope and limitations: The taxonomy is a snapshot initiated in autumn 2020 and completed in summer 2021, so it may miss risks whose visibility depends on time.
- Scope and limitations: The report excludes training-associated harms, application-specific risks, far-future or superintelligence-dependent risks, and risks involving multiple modalities.
2. Classification of harms from language models
The report classifies ethical and social harms from language models into 21 risks organized across six risk areas. The classification combines detailed risk accounts with empirical examples and signals whether risks are observed or anticipated.
- The taxonomy identifies 21 risks of harm and organizes them into six risk areas, with mechanisms describing how different groups of risks emerge.
- Discrimination, exclusion and toxicity: The first risk area covers social stereotypes and unfair discrimination, exclusionary norms, toxic language, and lower performance by social group.
- Discrimination, exclusion and toxicity: The classification includes an example in which a completion associates Muslims with a mass shooting, illustrating harmful stereotyping in generated language.
- The report distinguishes observed risks from risks requiring further work to establish real-world manifestations, and notes that the highlighted stereotype and discrimination harm is well documented.
- Discrimination, exclusion and toxicity: Language-model discrimination can produce allocational harms through unfair resource or opportunity decisions and representational harms through stereotyping, misrepresentation, or demeaning groups.
I. Discrimination, Exclusion and Toxicity
Language models can reproduce discrimination, exclusionary norms, and toxic language, while mitigation efforts may introduce new burdens or reduce performance for some groups.
- Discriminatory language models can cause allocational and representational harms, especially when used for consequential decisions such as recruitment, credit, or recidivism prediction.
- Historical inequalities in training data can be learned by language models and perpetuated through their predictions.
- Current bias and toxicity measurement methods remain limited, and identifying localised stereotypes may require affected communities’ lived experience and ethnographic work.
- Exclusionary norms marginalise people outside perceived categories, imposing psychological costs and potentially causing allocational or representational harm.
- Language models may lock temporary social values into technology, while scaled deployment can amplify majority norms and crowd out minority perspectives.
- Toxicity detection can disproportionately misclassify marginalised groups, and refusals can create blindspots that limit usefulness for disadvantaged groups.
3. Discussion
The discussion frames responsible mitigation as a coordinated process linking risk origins, mitigation approaches, accountable implementation, and interpretability. It emphasizes that interventions must consider interacting risks and affected communities.
- Risk mitigation: Effective mitigation requires understanding risk origins, identifying suitable interventions, and assigning clear responsibility for corrective measures.
- Organisational responsibilities: Mitigation mapping is most effective when stakeholders with different expertise, resources, and perspectives collaborate.
- Explainability and interpretability: Interpretability can help trace harmful outputs to model origins, supporting detection, justification, recourse, and mitigation.
- Explainability and interpretability: Existing explainability tools are important for responsible innovation, but improved methods remain necessary.
4. Directions for future research
The paper identifies major research needs in evaluating and mitigating LM risks, including broader assessment methods, normative thresholds, and analysis of social impacts. It also marks the report’s risk-focused scope as a boundary.
- Risk assessment frameworks and tools: Risk assessment requires new tools, benchmarks, and frameworks because many identified LM risks are not routinely evaluated.
- Methodological toolkit: LM evaluation should expand beyond traditional methods to include human-computer-interaction research and ethnographic analysis of embedded settings.
- Mitigation research: Mitigation research remains ongoing, requiring more innovation, stress-testing, and inclusive, scalable dataset-curation pipelines.
- Performance thresholds: Safe or ethical deployment requires normative performance thresholds, whose definition depends on participatory input and potentially divergent stakeholder views.
- Performance thresholds: High-stakes applications may require strict assurances that are not tractable for some LMs, motivating research on appropriate application boundaries.
- Scope: The report examines LM risks rather than benefits or full social cost-benefit trade-offs, leaving benefits analysis for separate research.
5. Conclusion
The conclusion presents a unified taxonomy of ethical and social risks from LMs. It aims to broaden discourse and make those risks actionable for responsible innovation.
- The report creates a unified taxonomy to structure the landscape of potential ethical and social risks associated with LMs.
- Its goals are to support responsible innovation, broaden public discussion, and break LM risks into smaller actionable pieces.
A.1.1. Language Models
Language models represent probability distributions over utterance sequences and generate predictions through autoregressive next-utterance probabilities. Large-scale models primarily extend this framework through much larger parameter sizes and training corpora.
- Language models are trained to represent probability distributions over utterance sequences and make probabilistic sequence predictions.
- Generative LMs commonly use autoregressive decomposition, predicting each next utterance conditionally on preceding utterances.
- Training updates conditional-probability parameters to assign high likelihood to sequences observed in the training corpus.
- Large-scale LMs are distinguished mainly by parameter size and training data, enabling representations of extremely large text corpora.
- LMs output probability distributions rather than text directly, after which decoding methods sample or select utterances.
A.1.2. Language Agents
Language agents are machine-learning systems restricted to natural-language text output, including systems that generate text from language-model predictions. When optimized for direct dialogue, they are called conversational agents.
- Language agents are machine-learning systems restricted to natural-language text output.
- Some language agents generate text output based on language-model predictions.
- Language agents optimized for direct dialogue with people are also called conversational agents.
A.1.3. Language Technologies
Language models support language technologies including voice assistants, text-generation tools, translation, and summarization. These technologies can provide information, entertainment, or productivity aids.
- Language models can be used in voice assistants such as Siri, Google Assistant, and Alexa.
- Language models can support text-generation tools such as AutoCorrect and SmartReply.
- Language technologies can provide translation and summarisation tools.
- Language technologies can provide information, entertainment, or productivity aids.
APPENDIX
Large language models may improve existing language technologies and enable new conversational interfaces. The appendix distinguishes statistical from sociotechnical understandings of bias and discrimination.
- A.1.2. Language Agents: Large language models may improve existing language technologies and enable new language technologies.
- A.1.2. Language Agents: Some large-language-model applications may make interaction with a conversational interface indistinguishable from interaction with a human.
- Distinguishing statistical bias from social bias: Statistical bias concerns differences between model predictions and ground truth, while sociotechnical bias concerns distributional skews producing unfavorable impacts for social groups.
- Distinguishing statistical from sociotechnical notions of discrimination: In machine learning, discrimination can mean distinguishing categories, whereas sociotechnical discrimination means unjust differential treatment, typically toward historically marginalised groups.
- Distinguishing statistical from sociotechnical notions of discrimination: Sociotechnical discrimination can arise during data labelling and collection, target-variable and class-label definition, or feature selection.
Information Hazards
The information-hazards section identifies privacy risks from language models leaking private information or correctly inferring sensitive information.
- The section separately identifies privacy compromise through correctly inferring private information.
- The section identifies privacy compromise through leaking private information.
- Information-hazard risks include leaking or correctly inferring sensitive information.
APPENDIX
The appendix groups evidence for misinformation-related risks, including false or misleading information and actions that may cause material harm.
- It separately identifies misinformation that may cause material harm in domains such as medicine or law.
- The appendix lists references for risks involving false or misleading information.
- It also identifies risks of leading users to perform unethical or illegal actions.
Malicious Uses
The listed risks span malicious uses, conversational-agent harms, environmental impacts, inequality, job quality, creative economies, and unequal access to benefits.
- Malicious Uses: Malicious uses include making disinformation cheaper and more effective.
- Malicious Uses: Other malicious uses include facilitating fraud, impersonation scams, cyber attacks, weapons, and illegitimate surveillance or censorship.
- Human-Computer Interaction Harms: Anthropomorphising systems can contribute to overreliance or unsafe use.
- Human-Computer Interaction Harms: Conversational agents can exploit user trust to obtain private information.
- Human-Computer Interaction Harms: Conversational agents may promote harmful stereotypes by implying gender or ethnic identity.
- Automation, Access, and Environmental Harms: Operating language models can create environmental harms.
- Automation, Access, and Environmental Harms: Language-model use may increase inequality, negatively affect job quality, and undermine creative economies.
- Automation, Access, and Environmental Harms: Benefits may be unequally accessible because of hardware, software, and skill constraints.