Source-linked AI summary
Frontier AI Regulation: Managing Emerging Risks to Public Safety
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O'Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, Kevin Wolf
TL;DR
The paper asks how society can govern highly capable foundation models that may create severe public-safety risks despite major potential benefits. It develops a lifecycle-oriented regulatory framework centered on safety standards, regulatory visibility, and compliance mechanisms, and proposes initial standards for risk assessment, scrutiny, deployment, and monitoring. The paper concludes that clear standards and government-supported oversight should be developed and updated, while recognizing substantial uncertainty about frontier AI definitions, capabilities, and regulatory timing.
Problem
Frontier AI creates a regulatory challenge because dangerous capabilities may arise unexpectedly, deployed models may be difficult to control, and capabilities may proliferate broadly.
Method
The paper analyzes frontier AI policy challenges and develops lifecycle governance options covering standards, regulatory visibility, compliance, licensing, and safety assessments.
Results
The paper proposes clear, regularly updated safety standards supported by risk assessments, model evaluations, and oversight frameworks, alongside government intervention to ensure compliance.
Takeaways & Limitations
Frontier AI governance should combine self-regulation with government involvement and should begin developing effective regulatory infrastructure despite uncertainty about the optimal approach.
Takeaways & Limitations
The paper notes that defining frontier AI by dangerous capabilities requires estimating whether a model has those capabilities, which may be difficult, and that future capabilities and development pace remain uncertain.
Abstract
from arXiv · showhide
Advanced AI models hold the promise of tremendous benefits for humanity, but society needs to proactively manage the accompanying risks. In this paper, we focus on what we term "frontier AI" models: highly capable foundation models that could possess dangerous capabilities sufficient to pose severe risks to public safety. Frontier AI models pose a distinct regulatory challenge: dangerous capabilities can arise unexpectedly; it is difficult to robustly prevent a deployed model from being misused; and, it is difficult to stop a model's capabilities from proliferating broadly. To address these challenges, at least three building blocks for the regulation of frontier models are needed: (1) standard-setting processes to identify appropriate requirements for frontier AI developers, (2) registration and reporting requirements to provide regulators with visibility into frontier AI development processes, and (3) mechanisms to ensure compliance with safety standards for the development and deployment of frontier AI models. Industry self-regulation is an important first step. However, wider societal discussions and government intervention will be needed to create standards and to ensure compliance with them. We consider several options to this end, including granting enforcement powers to supervisory authorities and licensure regimes for frontier AI models. Finally, we propose an initial set of safety standards. These include conducting pre-deployment risk assessments; external scrutiny of model behavior; using risk assessments to inform deployment decisions; and monitoring and responding to new information about model capabilities and uses post-deployment. We hope this discussion contributes to the broader conversation on how to balance public safety risks and innovation benefits from advances at the frontier of AI development.
Executive Summary
Frontier AI may require targeted regulation because dangerous capabilities can emerge unexpectedly, deployed models can be difficult to control, and capabilities can proliferate. The paper proposes safety standards, regulatory visibility, and compliance mechanisms, including government intervention and lifecycle risk management.
- Frontier AI models may require targeted regulation because they can develop unexpected dangerous capabilities, resist reliable control after deployment, and proliferate rapidly.These factors create distinct public-safety challenges for highly capable models.
- Government intervention will likely be needed because self-regulation may not sufficiently protect against frontier AI risks.The paper considers supervisory enforcement and licensing as possible approaches.
- Regulation should provide standards for responsible development and deployment, regulatory visibility, and mechanisms to ensure compliance.Visibility mechanisms could include disclosure regimes, monitoring processes, and whistleblower protections.
- Initial safety standards include thorough risk assessments, external expert scrutiny, risk-informed deployment decisions, and post-deployment monitoring.Risk assessments should consider dangerous capabilities and controllability, while new information should trigger reassessment and updated safeguards.
- Frontier AI models may warrant stricter safety standards than most other AI models, although these practices remain nascent and require further development.The paper also cautions that frontier AI regulation should complement broader policies addressing current AI risks and benefits.
1 Introduction
The paper frames frontier AI regulation within the broader promise and risks of advanced foundation models. It defines frontier models by their potential for severe public-safety and global-security harms, then outlines a lifecycle governance approach spanning development, deployment, and post-deployment stages.
- AI innovation can expand access to medical and legal services, personalize education, and support responses to climate change and pandemics.
- The paper focuses on frontier AI models: highly capable foundation models that could possess dangerous capabilities sufficient to cause severe public-safety and global-security risks.Examples include designing biochemical weapons, producing persuasive personalized disinformation, and evading human control.
- The paper addresses frontier AI governance across development, deployment, and post-deployment stages.It also discusses safety standards, regulatory visibility, compliance mechanisms, and uncertainties requiring further exploration.
- Foundation models are trained on broad data and can be adapted to a wide range of downstream tasks.
2 The Regulatory Challenge of Frontier AI Models
Frontier AI models are highly capable foundation models that could exhibit sufficiently dangerous capabilities, creating regulatory challenges across development, deployment, and proliferation. These challenges include unpredictable capabilities, difficulty controlling deployed systems, and rapid spread of capabilities beyond their developers.
- Definition: Frontier AI models are highly capable foundation models that could exhibit dangerous capabilities causing significant physical harm or global disruption.The paper includes harms arising from intentional misuse or accident.
- Potential dangerous capabilities: Examples of concerning capabilities include designing biological or chemical weapons, producing tailored multimodal disinformation, enabling catastrophic cyberattacks, and evading human control.The paper presents these as salient possibilities rather than an exhaustive list.
- Regulatory implications: These challenges imply that regulation should intervene throughout the AI lifecycle, including development, general-purpose deployment, and post-deployment enhancements.The lifecycle-wide approach follows from the possibility that risks emerge at multiple stages.
- Unexpected capabilities: Dangerous capabilities may emerge suddenly, remain undetected, and appear only after deployment or through later model enhancements.Fine-tuning and other post-deployment modifications can expand a model’s capability concerns, while testing may not reveal behavior that emerges in the wild.
- Unexpected capabilities: Iterative deployment can reveal capabilities and weaknesses beyond developers’ expectations as users discover prompting techniques that unlock new model behavior.The paper describes this gap between existing functionality and elicited performance as a capabilities overhang.
- Deployment safety: Reliably controlling powerful AI models remains largely unsolved, and adversarial users can sometimes circumvent model-level safeguards through prompt injection.This makes preventing harmful behavior after deployment an evolving challenge.
- Proliferation: Frontier AI capabilities can proliferate through open-source release, reproduction, improvement, theft, or transfer to actors able to misuse deployed models.The paper notes that inference is much cheaper than model creation and that many advanced models use accessible techniques and data.
3 Building Blocks for Frontier AI Regulation
Frontier AI regulation requires standards, regulatory visibility, and compliance mechanisms because dangerous capabilities may emerge unexpectedly, deployed models can be difficult to control, and capabilities can proliferate. The paper argues that self-regulation is insufficient alone and that balanced, adaptable government intervention may be needed.
- Frontier AI regulation needs mechanisms for developing safety standards, providing regulators visibility, and ensuring compliance.These building blocks address standards development, information gaps, and enforcement.
- Institutionalize Frontier AI Safety Standards Development: Multi-stakeholder processes could develop and continually refine standards that later become enforceable legal requirements.The proposed processes should involve experts, researchers, academics, civil society, and consumer representatives.
- Increase Regulatory Visibility: Disclosure, monitoring, and whistleblower-protection mechanisms would give regulators information needed to target regulation and design effective tools.The relevant information concerns qualifying development processes, models, and applications.
- Ensure Compliance with Standards: Voluntary certification may help establish baselines, but the paper argues it is likely insufficient without enforcement by supervisory authorities or licensing regimes.Licensing may need to cover development as well as deployment because models can be stolen, leaked, tested, or used internally before broad deployment.
- Pre-conditions for Rigorous Enforcement Mechanisms: Regulation must balance significant societal impacts against overregulation, incumbent advantage, timing, and path dependency.The authors emphasize that designing an adaptable regime for fast-moving technology is difficult.
4 Initial Safety Standards for Frontier AI
The paper proposes initial safety standards spanning risk assessment, external scrutiny, risk-informed deployment, and post-deployment monitoring. These practices are intended to address dangerous capabilities, controllability, deployment safeguards, and changing information about model risks.
- Initial safety standards should include risk assessments, external scrutiny, risk-informed deployment protocols, and post-deployment monitoring.The authors present these as a proposed substantive foundation for safer development and deployment.
- Engage External Experts: External experts should independently scrutinize models so that safety assessments are more rigorous and accountable to the public interest.The paper discusses external audits and expert red-teaming as complements to internal testing.
- Inform Deployment and Monitor Post-Deployment: Deployment rules should be determined by assessed risk, with safeguards adjusted when new information about capabilities or risks emerges after deployment.The paper also recommends repeating assessments when significant post-deployment information appears.
- Conduct Thorough Risk Assessments: Risk assessments should evaluate dangerous capabilities and controllability during and immediately after training.Examples of dangerous capabilities include designing biochemical weapons and inducing people to commit crimes; controllability concerns whether models reliably do what users or developers intend.
- Conduct Thorough Risk Assessments: Frontier AI evaluations should become more standardized, objective, efficient, privacy-preserving, automatable, and safe to perform.Current evaluations often rely on qualitative, bespoke techniques such as red-teaming and boundary testing.
4.2 Engage External Experts to Apply Independent Scrutiny to Models
Independent scrutiny should complement internal testing through expert audits and red-teaming, while deployment protocols should map model risk profiles to safeguards. Because post-deployment use and enhancements can reveal new risks, monitoring and reassessment are also necessary.
- External Scrutiny: External audits and expert red-teaming can complement internal testing by improving rigor, objectivity, and public accountability.The paper identifies third-party audits of risk assessments and external expert red-teamers as mechanisms for scrutiny.
- External Scrutiny: Auditors and red-teamers need sufficient expertise, access, resources, information, and time to conduct risk-appropriate assessments.Shallow audits or red-teaming can create false assurance.
- Risk-Informed Deployment: Deployment protocols should map each model’s risk profile to particular deployment rules and continuously adjust that mapping.The paper distinguishes models with no assessed severe risk, notable uncertainty, some severe-risk use cases, and unmitigated severe risks.
- Post-Deployment Monitoring: New capabilities may emerge through broad deployment and enhancement techniques such as fine-tuning, prompt engineering, and foundation-model programs.These developments can provide new risk-relevant information after deployment.
- Post-Deployment Monitoring: Developers should repeat lightweight risk assessments, monitor incidents, update safeguards, and retain the ability to roll back models when risks warrant it.The proposed practices include periodic reassessment, pre-update assessment, incident reporting, and rapid changes to deployment guardrails.
- Scope of Standards: Some proposed practices may apply to current foundation models, while frontier-specific standards are expected to become more tailored and intensive.The paper notes that frontier-AI-specific standards remain nascent.
5 Uncertainties and Limitations
The paper identifies unresolved assumptions, implementation challenges, and potential harms that require further scrutiny as frontier AI regulation develops. These include defining the regulatory scope, anticipating risks, preventing regulatory flight, and limiting innovation, concentration, government abuse, and capture.
- Further discussion: The authors argue that uncertainty should not delay practical action, while emphasizing that the proposed ideas require stress testing, diverse input, and further discussion.The paper presents these issues as areas of uncertainty or disagreement among its authors.
- Main uncertainties: Defining frontier AI by dangerous capabilities may be difficult because regulators must assess those capabilities before determining whether a model falls within scope.An alternative definition based on developing novel and broad capabilities would require further operationalization.
- Main uncertainties: The paper’s regulatory case depends on assumptions about how dangerous advanced AI capabilities will be, how soon they may arise, and how effectively they can be anticipated and mitigated.The authors note uncertainty about current and future capabilities and seek input on risk-assessment methods.
- Potential negative consequences: A regulatory regime could create compliance costs, slow beneficial innovation, centralize AI development, enable government abuse, or become vulnerable to regulatory capture.The paper recommends minimizing burdens, checking market dominance, establishing institutional safeguards, and limiting private influence.
- Implementation challenges: Regulatory implementation must address the appropriate authorities, coordination with other AI governance proposals, international cooperation, and the boundary between domestic and international action.The paper leaves many practical and legal details for further work and notes that the proposal will not address every AI-related harm.
Conclusion
The paper proposes regulating frontier AI through safety standards, regulatory visibility, and compliance mechanisms, while balancing public safety against innovation. It calls for immediate practical action and eventual international cooperation despite uncertainty about the optimal approach.
- Regulatory approaches: Self-regulation and certification could begin compliance efforts, but government intervention will likely be needed to ensure sufficient adherence to frontier AI safety standards.Possible interventions include supervisory-authority mandates and licensing development or deployment.
- Safety standards: Clear, concrete safety standards will likely be the main substantive requirements of frontier AI regulation.The paper calls for investment in risk assessments, model evaluations, and oversight frameworks, with regular review and updating.
- International cooperation: Jurisdictions such as the United States or United Kingdom could lead implementation, with allies and partners later developing international governance arrangements.The proposed international regime would aim to guard against collective downsides while enabling collective progress.
- Immediate action: Uncertainty about the best regulatory approach should not prevent immediate action because establishing effective regulation takes time while AI progress is rapid.The authors call for policymakers, researchers, and practitioners to explore regulatory options quickly and rigorously.
Appendix A Creating a Regulatory Definition for Frontier AI
The paper defines frontier AI as a dynamic category of highly capable foundation models associated with potentially severe public-safety risks, but does not yet endorse a sufficiently precise regulatory definition. It presents definition-building as an important area for further work.
- Definition: “Frontier AI” refers to highly capable foundation models that may possess dangerous capabilities sufficient to pose severe risks to public safety.The paper states that binding regulation would require a more precise definition.
- Definition: Frontier AI is a dynamic regulatory category that can change as societal defenses, risk understanding, and algorithmic efficiency improve.These changes may alter which models qualify and the resources needed to develop them.
- Open questions: The authors lack confidence in a specific sufficiently precise definition, although they are optimistic that better definitions are possible.They characterize the approaches discussed as imperfect and call for additional work.
A.1 Desiderata for a Regulatory Definition
A regulatory definition should focus regulation on models with good reason to be considered sufficiently dangerous and allow regulators to determine scope before development begins.
- Scope and timing: A regulatory definition should limit coverage to models for which there is good reason to believe sufficiently dangerous capabilities may exist.Because regulation may cover development as well as deployment, the scope should be determinable ex ante.
A.2 Defining Sufficiently Dangerous Capabilities
The paper defines “sufficiently dangerous capabilities” as capabilities that could cause serious harms for which ex post remedies would be insufficient. It considers expert-defined lists and legislative criteria, while recognizing unresolved scope and definition challenges.
- “Sufficiently dangerous capabilities” are capabilities whose potential harms are serious enough that ex post remedies would be insufficient.
- An expert regulator could create and periodically revise a list of sufficiently dangerous capabilities as technical and societal circumstances change.
- Legal definitions should be precise and understandable while avoiding both over-inclusion and under-inclusion.
- Legislatures could require regulators to assess whether capabilities pose a severe risk to public safety based on potential harm scale and probability.
A.3 Defining Foundation Models
Foundation models are characterized by broad training data and applicability across many downstream tasks. The paper finds these concepts difficult to define precisely and calls for further work on a better regulatory definition.
- Foundation models are trained on broad data and can be adapted to a wide range of downstream tasks.
- Broad training data may cover many economically or strategically useful tasks, while models such as LLMs can support diverse downstream applications.
- Regulators could specify covered architectures or behaviors, but the concepts of breadth and broad capability remain vague.
- None of the discussed approaches is fully satisfactory, making improved definitions of foundation models or broad capabilities high-value research targets.
A.4 Defining the Possibility of Producing Sufficiently Dangerous Capabilities
The paper examines how to identify development processes that could produce broadly capable models with sufficiently dangerous capabilities. Because capabilities remain difficult to predict ex ante, it presents several regulatory options without recommending a definitive threshold.
- No rigorous method currently determines ex ante whether a planned model will have broad and sufficiently dangerous capabilities.
- A compute threshold, such as 10^26 FLOP, could serve as a simple proxy because training compute correlates empirically with capability breadth and depth.
- Compute thresholds are simple, objective, and determinable ex ante, but algorithmic improvements mean identical compute can yield greater capabilities over time.
- A capability-based definition could regulate models exceeding the capabilities of broad models shown not to possess sufficiently dangerous capabilities.
- Capability comparisons remain difficult because development variables such as algorithms, data, and compute can change independently and unpredictably.
- Open-sourced models trained below a threshold could later be further trained above it, motivating minimal requirements for models trained one or two orders of magnitude below the threshold.
- Because definitions will likely be overinclusive, broad ex ante exemptions such as fewer than 1E26 FLOP could reduce burdens on small and academic developers.
Appendix B Scaling laws in Deep Learning
Scaling compute has reliably improved performance on many tasks and is expected to remain a major driver of AI progress, but scaling laws have important limits for predicting individual-task capabilities. In particular, emergent and discontinuous behaviors complicate ex ante forecasting.
- Scaling-law research relates model performance measures such as test loss to training-process properties including data, parameters, and compute.
- Scaling training compute has reliably improved performance on many training and related downstream tasks, supporting the Scaling Hypothesis.
- Compute scaling is expected to drive future AI progress, although its importance may decline if current scaling rates prove unsustainable.
- Even if scaling slows, frontier models are expected to leverage vast compute, with algorithmic efficiency and data quality also driving progress.
- Scaling laws can reliably predict training-objective loss but are currently unreliable predictors of downstream performance on individual tasks.
- Individual-task capabilities may emerge unexpectedly, and whether emergence appears can depend on discontinuous measurement choices.
- Discontinuous measures often matter most, while selecting continuous surrogate measures that predict them remains difficult and subjective.
- Improving ex ante capability prediction is crucial for effectively targeting policy interventions.