Source-linked AI summary
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
Alex Tamkin, Miles Brundage, Jack Clark, Deep Ganguli
TL;DR
Large language models raise open questions about their technical capabilities, limitations, and societal effects. The discussion summarizes these questions across capabilities, multimodality, deployment, bias, disinformation, and governance. It concludes that broad capabilities coexist with unresolved challenges in safe deployment, bias mitigation, and alignment with human values.
Problem
The discussion addresses limited understanding of large language models' technical capabilities and limitations and the societal effects of their widespread use.
Method
The authors provide a detailed discussion summary organized around technical and societal themes, followed by potential future research directions.
Results
The discussion identifies broad capabilities and growing multimodal potential alongside unresolved challenges in scoping uses, safe deployment, disinformation, bias, and alignment.
Takeaways & Limitations
The discussion supports continued research and careful governance of large language models, including improved evaluation, deployment practices, and bias mitigation.
Takeaways & Limitations
GPT-3's unusually broad capability surface makes it difficult to anticipate all potential behaviors and ensure safe use.
Abstract
from arXiv · showhide
On October 14th, 2020, researchers from OpenAI, the Stanford Institute for Human-Centered Artificial Intelligence, and other universities convened to discuss open research questions surrounding GPT-3, the largest publicly-disclosed dense language model at the time. The meeting took place under Chatham House Rules. Discussants came from a variety of research backgrounds including computer science, linguistics, philosophy, political science, communications, cyber policy, and more. Broadly, the discussion centered around two main questions: 1) What are the technical capabilities and limitations of large language models? 2) What are the societal effects of widespread use of large language models? Here, we provide a detailed summary of the discussion organized by the two themes above.
Introduction
The discussion examined both the technical capabilities and limitations of large language models and the societal effects of their widespread use. The document summarizes these themes and identifies potential future research directions.
- The meeting brought together researchers from computer science, linguistics, philosophy, political science, communications, and cyber policy.
- The technical discussion addressed scale-driven capabilities, language understanding, multimodal training, and alignment with human values.
- The societal discussion addressed the difficulty of anticipating uses and misuses, deployment challenges, disinformation, bias, and labor-market effects.
- The authors provide a detailed summary organized around the two themes and conclude with potential future research directions.
1 Technical Capabilities and Limitations
Participants discussed GPT-3's emergent capabilities, the unresolved meaning of language understanding, the promise of multimodal training, and the challenge of aligning models with human values. They emphasized that impressive performance does not settle questions about robustness, causality, or understanding.
- Scale: 175 billion parameters and 570 gigabytes of training text enabled GPT-3 to learn novel tasks from in-context examples, extending GPT-2's zero-shot generalization.GPT-2 had 1.5 billion parameters and 40 gigabytes of training text.
- Scale: Scaling produced capability gains that participants described as unusually stable and predictable, potentially supporting stronger few-shot learning in larger models.
- Scale: Large-model development requires diverse teams to build infrastructure, develop algorithms, interrogate capabilities, and red-team for bias, misuse, and safety concerns.
- Understanding: Participants disagreed about whether GPT-3 understands language, invoking intentionality, real-world responsiveness, adversarial robustness, causal reasoning, and the possibility of shortcut features.
- Understanding: Some participants rejected a binary or singular notion of understanding, while others argued that understanding may be irrelevant to successful task performance.
- Multimodality: Multimodal models may become more prevalent and enable diverse capabilities, although GPT-3 already incorporated prose, structured tables, and code in its training data.
- Alignment: Alignment remains challenging because humans value factual accuracy and robustness differently from superficial token-level correctness.
2 Effects of Widespread Use
The discussion examined how GPT-3’s broad capabilities create deployment, misuse, bias, and labor-market challenges. Participants considered technical and governance measures, while emphasizing that access controls and mitigation strategies involve unresolved tradeoffs.
- Deployment and governance: GPT-3’s broad capability surface makes it difficult to anticipate all potential uses and ensure safety, while controlled API access can constrain use more easily than open sourcing.Open questions include who receives access, why, and how to support large-scale red-teaming.
- Governance and societal evaluation: Participants discussed increasing academic computing resources, disclosure requirements for AI-generated text, and metrics for evaluating whether language models have societally beneficial effects.They agreed that measuring societal benefit is challenging but important.
- Misuse and disinformation: Models like GPT-3 can generate false, misleading, or propagandistic essays, tweets, and news stories de novo, raising concerns about deliberate disinformation misuse.Future systems might produce persuasive text on arbitrary topics, potentially forcing platforms to rely on metadata and authenticated media.
- Bias: GPT-3 exhibits racial, gender, and religious biases, but participants emphasized that defining mitigation universally is difficult because appropriate language use is contextual.The challenge is addressing harmful biases according to normative or legal criteria rather than eliminating all bias.
- Bias: Proposed bias-mitigation approaches included training-data changes, filters, fine-tuning, data tagging, fact-aware training, human feedback, prompt design, bias tests, and scaled red-teaming.These approaches target different stages of model development and deployment.
- Bias: No mitigation approach was considered a panacea: human-feedback steering raises questions about labeler selection, while filters can undermine the agency of groups they aim to protect.The discussion therefore treated mitigation as involving normative choices as well as technical interventions.
- Labor-market effects: Language-model automation raises questions about which text-related jobs should be automated because such work ranges from enjoyable creative tasks to traumatizing or alienating content moderation.Participants connected automation decisions to the varied desirability of existing jobs.
3 Future Research Directions
The proposed research agenda spans technical questions about scaling, reasoning, multimodal learning, and uncertainty, alongside societal questions about misuse, access, safety, guardrails, and bias. Together, these directions address both model capabilities and the governance of their deployment.
- Technical capabilities and scaling: Future work should explain why language models improve with scale and determine whether scaling can become more efficient.This agenda links understanding scaling behavior with reducing the resources required for capability growth.
- Technical capabilities and scaling: Researchers should test whether scaling yields causal reasoning, symbolic manipulation, commonsense understanding, and robustness, or whether different techniques are necessary.The question concerns the limits of scaling rather than assuming continued capability growth.
- Technical capabilities and uncertainty: Future models should better recognize their capability limits by asking for help or clarification, or abstaining when uncertain.This direction targets behavior under uncertainty.
- Multimodal learning: New architectures and algorithms should enable efficient learning from diverse, multimodal data beyond text.The proposed direction expands learning beyond text-only inputs.
- Alignment: Research should characterize the opportunities and tradeoffs of steering large language models toward greater alignment with human values.The question treats alignment as involving competing considerations rather than a single optimization target.
- Access and safety: Future work should determine how model access can balance security, replicability, and fairness, and develop tests for safety in particular contexts.The agenda links access allocation with context-specific qualification as safe or unsafe.
- Institutions and bias: The agenda includes positioning academia to develop industrial guardrails and fostering cross-disciplinary collaboration on biases in datasets and model representations.Both directions address institutional capacity and expertise needed to manage large-model risks.
- Misuse and threat assessment: The threat landscape should be compared across profit-driven spam and state-based disinformation, including the cost-effectiveness and skill requirements of different misuse strategies.These questions explicitly compare malicious uses with alternative ways of achieving the same goals.