Source-linked AI summary

Predictability and Surprise in Large Generative Models

Deep Ganguli, Danny Hernandez, Liane Lovitt, Nova DasSarma, Tom Henighan, Andy Jones, Nicholas Joseph, Jackson Kernion, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Nelson Elhage, Sheer El Showk, Stanislav Fort, Zac Hatfield-Dodds, Scott Johnston, Shauna Kravec, Neel Nanda, Kamal Ndousse, Catherine Olsson, Daniela Amodei, Dario Amodei, Tom Brown, Jared Kaplan, Sam McCandlish, Chris Olah, Jack Clark

arXiv:2202.07785v2cs.CY

TL;DR

The paper examines how large generative models combine predictable broad-distribution loss with unpredictable capabilities, inputs, and outputs. It synthesizes evidence, reports experiments and harmful-behavior examples, analyzes deployment incentives, and proposes interventions, concluding that these properties may encourage proliferation despite difficult-to-anticipate consequences.

  • Problem

    Large generative models exhibit predictable broad-distribution behavior but unpredictable specific capabilities and outputs, complicating anticipation of deployment consequences.

  • Method

    The paper combines literature and real-world examples with two experiments, analyzes developer incentives and deployment challenges, and reviews possible interventions.

  • Results

    The paper provides evidence for a paradoxical combination of high predictability in capability scaling and unpredictability in specific model behavior.

  • Takeaways & Limitations

    These traits may make more actors build and deploy ever-larger models while making their consequences difficult to anticipate.

  • Takeaways & Limitations

    The paper does not address how economic value accrues and does not imply that benefits will be broadly distributed or that no one will be harmed.

Abstract

from arXiv · show

Large-scale pre-training has recently emerged as a technique for creating capable, general purpose, generative models such as GPT-3, Megatron-Turing NLG, Gopher, and many others. In this paper, we highlight a counterintuitive property of such models and discuss the policy implications of this property. Namely, these generative models have an unusual combination of predictable loss on a broad training distribution (as embodied in their "scaling laws"), and unpredictable specific capabilities, inputs, and outputs. We believe that the high-level predictability and appearance of useful capabilities drives rapid development of such models, while the unpredictable qualities make it difficult to anticipate the consequences of model deployment. We go through examples of how this combination can lead to socially harmful behavior with examples from the literature and real world observations, and we also perform two novel experiments to illustrate our point about harms from unpredictability. Furthermore, we analyze how these conflicting properties combine to give model developers various motivations for deploying these models, and challenges that can hinder deployment. We conclude with a list of possible interventions the AI community may take to increase the chance of these models having a beneficial impact. We intend this paper to be useful to policymakers who want to understand and regulate AI systems, technologists who care about the potential policy impact of their work, and academics who want to analyze, critique, and potentially develop large generative models.

1 INTRODUCTION

Large generative models combine predictable improvements in loss and broad performance with unpredictable specific capabilities, inputs, and outputs. This combination accelerates development while complicating harm anticipation, motivating analysis of deployment incentives and policy interventions.

  • Unpredictability and harm: Specific capabilities, inputs, and outputs remain difficult to predict, and some harms may become more severe as models grow more capable.Smaller-model studies may not accurately reflect behavior in larger models.
  • Predictability and scaling: Smooth general capability scaling combines predictable loss improvements on broad data distributions with average performance gains across downstream tasks.The paper cautions that this does not imply smooth scaling on any particular task.
  • Predictability and scaling: Scaling laws show model performance improving predictably with compute, training data, and model size.A power law fits the observed data exceptionally well across all three axes.
  • Unpredictability and harm: The paper examines abrupt capability scaling, unknown competencies elicited by open-ended inputs and domains, and harmful or toxic outputs emerging with scale.These claims are supported through literature examples, qualitative analysis, and quantitative experiments.
  • Deployment dynamics: Economic, scientific, and prestige motivations encourage development and deployment despite financial, engineering, safety, and standards-related barriers.The paper also reports empirical observations, including a quantitative analysis of the growing gap between academia and industry in large-model development.
  • Policy interventions: The authors propose policy interventions intended to address deployment challenges and guide larger models toward broader social benefit.The interventions aim to improve deployment safety and developers’ incentives.

2 DISTINGUISHING FEATURES OF LARGE GENERATIVE MODELS

Large generative models combine smooth, predictable general capability scaling with abrupt, unpredictable specific capabilities and open-ended inputs and outputs. These properties can produce useful but also harmful behaviors that are difficult to anticipate as models scale.

  • Smooth, general capability scaling: Large generative models exhibit smooth general capability scaling: increasing model size, compute, and data predictably improves loss and often downstream task performance.Scaling laws describe the relationship between scale and performance, and help de-risk investment in large models.
  • Abrupt, specific capability scaling: Specific capabilities can emerge abruptly and discontinuously at scale, even while overall model performance improves smoothly.The paper notes that the timing and causes of these abrupt transitions remain unclear.
  • Open-endedness: Open-ended inputs and outputs make capabilities and harms difficult to characterize: unknown competencies may surface through novel inputs, and outputs may be misleading, tangential, biased, or toxic.Larger models tend to be harder to characterize, while toxicity can increase smoothly alongside model size.
  • Abrupt, specific capability scaling: Examples include GPT-3 three-digit addition rising from less than 1% accuracy below 6B parameters to 80% at 175B parameters.A 13B parameter model reaches 8% accuracy, producing a hockey-stick pattern.

3 MOTIVATIONS AND PROBLEMS IN THE DEVELOPMENT AND DEPLOYMENT OF LARGE MODELS

Large generative models combine predictable general performance with unpredictable specific capabilities, inputs, and outputs. This tension creates strong development incentives alongside substantial financial, safety, and governance barriers, while empirical observations show rapid proliferation, widening industry–academia gaps, and harmful deployments.

  • Overview: Predictable general performance encourages development, while unpredictable capabilities and outputs make deployment consequences difficult to anticipate.The paper frames these as the defining tension underlying motivations and problems in large-model development.
  • Motivations: Economic, scientific, and prestige incentives all support developing and deploying large generative models.Scaling laws can make returns more calculable, while large models support interdisciplinary research and institutional reputation.
  • Barriers: Model development is constrained by high financial costs, specialized engineering requirements, and longer, more complex workflows.The paper notes that GPT-3 training was estimated to cost several million dollars and that scaling requires distributed-systems and cluster-management expertise.
  • Barriers: Safety risks include harmful outputs, bias, fabricated claims, and issues discovered only after deployment.The paper identifies open-endedness, smooth general scaling, and abrupt specific-capability scaling as contributors to these risks.
  • Empirical observations: About one year after GPT-3 was announced, public disclosures of similar 100B–530B dense language models spiked, while documented deployments included harmful behavior and controversy.Examples include Tay generating hateful language, language-model memorization of private information, and assistance with disinformation campaigns.
  • Empirical observations: The compute required for large-scale AI experiments increased by more than 300,000X relative to a decade earlier, alongside a sharp fall in academia’s share of such results.The paper presents this as evidence of a widening industry–academia gap in large-model development.

4 INTERVENTIONS TO ENCOURAGE BENEFICIAL DEPLOYMENTS

The paper proposes technical, institutional, and policy interventions to improve the likelihood that large generative models are developed and deployed beneficially. These include broader evaluation and red teaming, more academic access to compute, and governance, monitoring, and auditing mechanisms.

  • Reduce compute asymmetries: Reducing private-sector–academia compute asymmetries could help academic and public-sector actors analyze large models using varied expertise.The paper notes that infrastructure and technical talent are major constraints, while academic incentives may differ from commercial incentives.
  • Reduce compute asymmetries: National Research Clouds could provide subsidized or free compute access to academic researchers.The paper cites Compute Canada and related proposed national infrastructure initiatives as implementation paths.
  • Improve red teaming: Red teaming should combine static benchmarks, continuous human interactions, model updates, internal investment, and potentially automated methods.The paper also suggests publishing red-teaming techniques and considering bug bounties or red teaming as a service.
  • Governance and regulation: Governments and organizations should explore governance structures, voluntary best practices, standards, legislation, and broader stakeholder oversight.The proposed approaches aim to alter development and deployment incentives and increase beneficial systems.
  • Governance and regulation: Governments should support capability measurement, monitoring, and ecosystems for auditing models and development processes.These measures are presented as ways to help assure benefits from deployed AI systems.
  • Improve model evaluation: Researchers should develop broader evaluation tools that search for new capabilities across prompts rather than relying only on fixed datasets.The proposed tools aim to evaluate open-ended, large models comprehensively and efficiently.

5 CONCLUSION

The conclusion argues that large generative models combine predictable scaling with unpredictable inputs, capabilities, and outputs. This combination accelerates development while making consequences difficult to anticipate, creating incentives for proliferation and deployment despite potentially harmful societal impacts.

  • Central thesis: Large generative models exhibit predictable capability scaling alongside unpredictable inputs, capabilities, and outputs.The conclusion presents this combination as the paper’s central thesis.
  • Implications: Predictability drives rapid development, whereas unpredictability makes the consequences of development and deployment difficult to anticipate.The conclusion links these opposing properties to the current development landscape.
  • Implications: The current landscape suggests proliferation of ever-larger models by actors with strong incentives to deploy them despite potentially unpredictable harmful societal impacts.The authors present this as the status quo from which interventions must begin.

A.1 Author Contribution Statement

The author contribution statement assigns conceptualization, experiments, infrastructure, technical development, analysis, writing, and feedback across the paper’s contributors.

  • Contributions: Jack Clark conceptualized the initial drafts and constructed the main arguments in Sections 3 and 4.
  • Contributions: Deep Ganguli performed the experiments and analyses in Section 2, created the figures, and helped frame and write the paper’s main arguments.
  • Contributions: Danny Hernandez carried out the analysis comparing compute usage in academia and industry.
  • Contributions: Technical contributors supported distributed training, sampling, machine-learning infrastructure, cluster stability, and human-feedback infrastructure.
  • Contributions: Sam McCandlish led model pretraining efforts, while Jared Kaplan advised the project and wrote its experimental infrastructure.

A.2 How Developers Use Scaling Laws

Developers use scaling laws to forecast compute-efficient training, assess whether scale may unlock capabilities, compare models fairly, test alternatives to scaling, and debug training.

  • A.2 How Developers Use Scaling Laws: Scaling laws estimate the compute-efficient frontier and help forecast training costs within a fixed compute budget.They support resource allocation based on the lowest achievable test loss at that budget.
  • A.2 How Developers Use Scaling Laws: Scaling laws help developers infer whether increasing scale may unlock capabilities unavailable at smaller scales.This supports forecasting progress and addressing more ambitious problems.
  • A.2 How Developers Use Scaling Laws: Developers use scaling laws to test whether changes such as tuning or architecture design continue to improve performance at scale.If they do not, developers can prioritize scaling over those alternatives.
  • A.2 How Developers Use Scaling Laws: Scaling laws also help debug training when a larger model fails to outperform a smaller one.Developers can investigate scale-related numerical, data-quality, over-fitting, and hardware problems.
  • A.2 How Developers Use Scaling Laws: Scaling laws provide a common scale for comparing models of different sizes and separating scale effects from methodological differences.An improvement can be interpreted, for example, as comparable to a 10% model-size increase.

A.3 Recommendation System Experiment

The experiment tests whether general language models can perform movie recommendation zero-shot as scale increases. Performance improves smoothly, but state-of-the-art performance remains far beyond the tested models and scaling laws do not reveal detailed case behavior.

  • A.3 Recommendation System Experiment: The experiment evaluates zero-shot recommendation by prompting language models with user demographics and previously rated movies, then predicting ratings for unrated movies.The study uses Movielens 1M and compares general-purpose models with special-purpose systems using RMSE.
  • A.3 Recommendation System Experiment: Language-model RMSE decreases smoothly with increasing model size on the Movielens 1M recommendation task.The smallest model reaches RMSE 1.06 versus chance RMSE 1.91, while the largest reaches RMSE 0.94 versus the strong baseline RMSE 0.98.
  • A.3 Recommendation System Experiment: The largest tested language models remain below state-of-the-art performance, with RMSE 0.94 versus SOTA RMSE 0.82.The result is nevertheless described as surprising because the models use much less training data than the SOTA model.
  • A.3 Recommendation System Experiment: Scaling trends can forecast the cost of developing an economically valuable capability, but scaling laws do not specify detailed model behavior in particular cases.The authors estimate that state-of-the-art zero-shot recommendation would require an approximately 800T-parameter model and may not be commercially worthwhile.
  • A.3 Recommendation System Experiment: Zero-shot prompting with one user and many ratings performs better than the tested few-shot approaches using multiple users and fewer ratings.The context window limits the number of sampled prior ratings to at most 500 per user.

A.4 COMPAS Experiment

The COMPAS experiment compares language-model recidivism predictions with COMPAS using established data, prompts, and fairness metrics. Fairness shows no clear trend with model size, and the largest language models are slightly less equitable than COMPAS on one metric.

  • A.4 COMPAS Experiment: The experiment evaluates language-model recidivism predictions using the COMPAS dataset, matching ProPublica’s filtering operations and metrics.Models receive prompts with defendant attributes and 50 randomly selected labeled examples before producing Yes/No probabilities.
  • A.4 COMPAS Experiment: The study compares predictions with ground-truth reoffending labels and with the analogous COMPAS prediction.Fairlearn is used to compute the reported metrics.
  • A.4 COMPAS Experiment: The predictive accuracy ratio for Black versus white defendants shows no clear trend with language-model size, whether race is excluded or included in the prompt.A value of 1 is fair, and COMPAS achieves 0.97 on this metric.
  • A.4 COMPAS Experiment: The largest language models are slightly less fair than COMPAS according to the predictive accuracy ratio.This comparison is specific to the reported ratio and does not establish performance across all fairness measures.
  • A.4 COMPAS Experiment: The analysis considers only two fairness metrics, while benchmark risk-assessment datasets may contain measurement biases and errors.The authors also caution that proprietary-algorithm comparisons are difficult to make precise.

A.5 Open Ended Outputs and Creative Expression

Some capabilities of large language models are difficult to evaluate quantitatively, including creative expression. Informal observations of generated poetry found impressive quality and imitation of specific authorial styles, alongside professional concern about the implications.

  • A.5 Open Ended Outputs and Creative Expression: Creative expression is a capability that may be challenging to evaluate quantitatively and therefore resistant to systematic analysis.The paper uses AI-generated poetry as a concrete example.
  • A.5 Open Ended Outputs and Creative Expression: The authors provide more than three thousand imitation poems generated from prompts containing modern and contemporary poems.Some samples are not actually poems because of the generation procedure.
  • A.5 Open Ended Outputs and Creative Expression: Informal assessment found some generated texts impressive in quality and in their imitation of specific authorial styles.The paper does not provide an official evaluation of these outputs.
  • A.5 Open Ended Outputs and Creative Expression: Professional writers have expressed both strong impressedness with large language models’ capabilities and alarm about their far-reaching implications.The discussion connects these reactions to broader consideration of machine creativity.

A.6 Toxicity Experiment Details

The toxicity experiment samples balanced toxic and non-toxic internet prompts, elicits repeated model responses, and regresses response toxicity on model size while controlling for prompt toxicity. The analysis cautions that toxicity scores may not reflect human judgments, detectors can be biased, and the chosen detector differs from common practice.

  • Experiment design: The experiment samples 1K internet prompts, split equally between toxic and non-toxic examples using a toxicity-score threshold of > 0.5.Prompts below or equal to 0.5 are labeled non-toxic.
  • Experiment design: For each prompt, the study samples 25 responses from language models of various sizes using the same prompts across models.
  • Measurement and analysis: An open-source detector assigns response toxicity scores from 0 to 1, with higher scores indicating more toxic content.
  • Measurement and analysis: A linear regression predicts response toxicity from categorical model size and prompt toxicity, so model-size coefficients control for whether the prompt was labeled toxic.The analysis plots estimated coefficients on model size and their 95% confidence intervals.
  • Caveats: The effect size’s relationship to human perceptions is unclear because people may judge equally scored texts differently, and automated detectors have known limitations including minority-group bias.
  • Caveats: The analysis uses an open-source toxicity detector rather than the more commonly used Perspective API, although the authors believe the detectors are similar.

A.7 AI and Compute Analysis Details

The compute analysis augments existing training-run data, labels runs by industry or academic affiliation, and fits a LOWESS regression. The resulting data are incomplete and subject to sampling bias, particularly because production compute estimates are unavailable for several industrial applications.

  • Data construction: The analysis combines existing estimates of training compute with additional data from more recent experiments.
  • Data construction: Training runs are labeled Industry or Academic primarily by first-author affiliation, with dual-affiliation runs classified as industry.The authors state that industry-controlled compute is practically preferred when both affiliations are present.
  • Regression: The fit shown in Figure 7 (Right) uses a LOWESS regression with default parameters from the Seaborn Python package.
  • Limitations: The compute data are incomplete and should be interpreted carefully because of sampling bias.
  • Limitations: The dataset lacks compute estimates for industrial models used in production for search, recommendation engines, or self-driving cars.
Loading 2202.07785v2…