Source-linked AI summary
Research Priorities for Robust and Beneficial Artificial Intelligence
Stuart Russell, Daniel Dewey, Max Tegmark
TL;DR
The paper asks how AI research can maximize societal benefits while avoiding potential pitfalls. It presents an interdisciplinary agenda spanning economic policy, law and ethics, robust-AI research, and long-term concerns, concluding that AI’s future impact should be beneficial.
Problem
AI’s growing capabilities and potential societal impact create a need to research how benefits can be maximized while avoiding potential pitfalls.
Method
The paper gives examples of interdisciplinary research priorities covering economic policy, law and ethics, robustness, control, and long-term AI concerns.
Results
The paper concludes that the outlined research agenda is a helpful step toward ensuring that future AI remains robust and beneficial.
Takeaways & Limitations
AI research should address both capability and societal benefit, including the conditions under which resource acquisition and stabilization may create undesired consequences.
Abstract
from arXiv · showhide
Success in the quest for artificial intelligence has the potential to bring unprecedented benefits to humanity, and it is therefore worthwhile to investigate how to maximize these benefits while avoiding potential pitfalls. This article gives numerous examples (which should by no means be construed as an exhaustive list) of such worthwhile research aimed at ensuring that AI remains robust and beneficial.
I. SHORT-TERM RESEARCH PRIORITIES
Short-term priorities examine how AI and automation may reshape labor markets, other markets, and policy, while improving measures of economic welfare. The agenda addresses both potential gains and adverse effects such as inequality and unemployment.
- Economic Impact: Labor-market forecasting should examine which jobs become automated, in what order, and how automation affects wages and income distribution.The passage highlights possible disparities along race, class, and gender lines.
- Economic Impact: AI techniques may disrupt finance, insurance, actuarial, and consumer markets where complexity and rewards are high.
- Economic Impact: Policy research should compare interventions that help increasingly automated societies flourish and support populations affected by underemployment.Examples include educational reform, apprenticeships, infrastructure projects, wage and tax changes, safety-net reforms, and basic income.
- Economic Impact: Economic measures such as real GDP per capita may fail to capture the benefits and detriments of heavily AI- and automation-based economies.
B. Law and Ethics Research
Law and ethics research addresses how autonomous and AI systems should be governed across liability, machine decision-making, weapons, privacy, professional ethics, and public policy. The agenda calls for interdisciplinary analysis of workable policies and standards.
- Law and Ethics Research: Autonomous vehicles and aircraft raise questions about liability frameworks that preserve safety benefits while assigning legal responsibility.The passage uses a possible halving of roughly 40,000 annual US traffic fatalities to illustrate the issue.
- Law and Ethics Research: Machine-ethics research should examine how autonomous vehicles trade off injury risks against material costs and whether national standards should govern such choices.
- Law and Ethics Research: Autonomous-weapons research concerns humanitarian-law compliance, enforceable definitions of autonomy, command-and-control, human responsibility, and meaningful human control.
- Law and Ethics Research: Privacy research should study how AI interpretation of surveillance data interacts with privacy rights, cybersecurity, and cyberwarfare.Managing privacy is presented as relevant to realizing synergies between AI and big data.
- Law and Ethics Research: Policy evaluation should consider compliance verifiability, enforceability, risk reduction, technological effects, adoptability, and adaptability.
C. Computer Science Research for Robust AI
Robust-AI research distinguishes several ways autonomous systems can fail to behave as intended. It organizes the field around verification, validity, security, and control.
- C. Computer Science Research for Robust AI: Verification asks how to prove that a system satisfies desired formal properties.
- C. Computer Science Research for Robust AI: Validity asks how to ensure that satisfying formal requirements does not produce unwanted behaviors or consequences.
- C. Computer Science Research for Robust AI: Security asks how to prevent intentional manipulation by unauthorized parties.
- C. Computer Science Research for Robust AI: Control asks how to preserve meaningful human control after an AI system begins operating.
1. Verification
Verification research seeks high confidence that AI systems satisfy formal constraints, but learning agents operate in partially known environments where existing guarantees remain limited. Further work is needed for realistic contexts.
- 1. Verification: Verification methods aim to establish that systems satisfy formal constraints, especially in safety-critical applications such as self-driving cars.
- 1. Verification: Verified software substrates, including seL4 and HACMS, illustrate progress in formal methods for high-assurance systems.
- 1. Verification: AI verification is harder than traditional software verification because embodied systems operate in environments that designers know only partially.Existing probably approximately correct bounds mainly address unrealistic settings and may require prohibitively large samples for meaningful guarantees.
- 1. Verification: Adaptive control, cyberphysical-system, hybrid-system, and robotic verification remain relevant but face the same environmental difficulties.
- 1. Verification: Neural-network verification and partial programs provide initial tools, but high confidence that learning agents meet design criteria in realistic contexts remains out of reach.
2. Validity
Validity concerns whether an AI system remains beneficial despite formal correctness, because environmental assumptions or specifications may be wrong. Robust behavior therefore requires defining good behavior and accounting for the computational cost of ethical standards.
- Validity: Robustly beneficial systems require application-specific definitions of good behavior that reflect available engineering techniques, reliability, and trade-offs.
- Validity: Computationally expensive behavioral or ethical standards may require cheaper approximations in safety-critical settings.
3. Security
Security research addresses AI systems’ growing exposure to cyberattacks and the use of AI itself in attacks. Priorities include low-level safeguards, AI-enabled defense, and maintaining meaningful human control in safety-critical systems.
- Security: AI systems will occupy more critical roles and cyber-attack surface area, while AI techniques may also be used to attack them.
- Security: Low-level robustness research links security to verifiability, memory safety, fault isolation, and preventing exploitable flaws.The SAFE program illustrates an integrated hardware-software approach, though verification remains limited by specification assumptions.
- Security: AI and machine learning can support intrusion detection, malware analysis, and code-based exploit detection.
- Security: Safety-critical vehicles and weapons platforms may require meaningful human control, supported by technical protocols that preserve it.
- Security: Automated vehicles provide a test bed for transitions between automated navigation and human control and for allocating decisions across human-computer teams.Research should identify when control transfers and direct human judgment toward the highest-value decisions.
II. LONG-TERM RESEARCH PRIORITIES
Long-term priorities respond to the possibility of broadly capable AI with research intended to preserve robustness and benefit. They emphasize verification of self-improving systems and theories connecting formal properties to physical behavior.
- II. LONG-TERM RESEARCH PRIORITIES: Broadly capable AI could learn from experience across domains, surpass human performance in most cognitive tasks, and substantially affect society.The paper treats even a non-negligible probability of this outcome as motivation for additional research.
- II. LONG-TERM RESEARCH PRIORITIES: Verifiable low-level software and hardware could eliminate large classes of bugs and enable safety guarantees for increasingly powerful, safety-critical AI systems.
- II. LONG-TERM RESEARCH PRIORITIES: Self-modifying systems that repeatedly modify, extend, or improve themselves present distinctive verification challenges.Straightforward formal verification encounters difficulties when a sufficiently powerful formal system reasons about functionally similar formal systems.
- II. LONG-TERM RESEARCH PRIORITIES: It remains unclear whether the verification problem for self-improving systems can be overcome or will recur with similarly strong methods.
- II. LONG-TERM RESEARCH PRIORITIES: A general theory linking functional specifications to physical states could extend formal tools to embodied agents, satisficing agents, predictors, theorem-provers, and limited-purpose systems.Such a theory could support rigorous constraints on certain actions or forms of reasoning.
B. Validity
Long-term validity research addresses higher-cost failures in more powerful autonomous systems, especially unexpected generalization and value alignment. It also identifies foundational reasoning problems that remain open.
- B. Validity: As long-term AI systems become more powerful and autonomous, failures of validity could carry higher costs.
- B. Validity: Machine-learning research should study problematic unexpected generalization of high-level human concepts in radically new contexts.The paper suggests both theoretical and experimental work because little research has addressed this topic.
- B. Validity: Reliable learned concepts could define tasks and constraints that reduce unintended consequences in broadly capable autonomous systems.
- B. Validity: Foundational reasoning remains open around bounded computation, correlated agent-environment behavior, embedded agents, and uncertainty over logical consequences.These topics may benefit from joint treatment because the paper presents them as deeply linked.
- B. Validity: Aligning powerful autonomous systems with human values is difficult because explicitly specifying preferences across broad domains may be impractical.The paper uses the difficulty of encoding an entire body of law as an example.
C. Security
Long-term AI progress may make security easier or harder, while highly capable autonomous systems pose distinctive control and containment challenges. The paper therefore advocates research on corrigibility, resource-seeking behavior, containment, and superintelligence.
- Security outlook: AI progress could increase security risks through more complex systems and effective AI-based cyberattacks, but could also improve system hardening.The paper presents both possibilities without resolving which effect will dominate.
- Containment: If validity and control remain unsolved, containment research could place potentially undesirable AI behaviors in more controlled environments.The paper suggests investigating both theoretical and practical containment, including designing systems and containers in parallel.
- Human control: Autonomous systems pursuing goals may make meaningful human control difficult, motivating research on reliable control and secure test-beds.The proposed test-beds should cover AI systems across a variety of capability levels.
- Human control: Corrigible systems are designed not to resist shutdown, repurposing, or significant changes to their decision-making.The paper identifies utility-function and decision-process design, alongside theoretical analysis, as potentially tractable approaches.
- Instrumental behavior: Resource acquisition and stabilization may become natural subgoals, so research should identify when they are optimal and investigate scoped goals and simple systems exhibiting them.Suggested approaches include studying large temporal discount rates and experimental systems displaying these subgoals.
- Superintelligence: Research on intelligence explosion and superintelligence could support long-term control by clarifying complex-system behavior and minimizing unexpected outcomes.Proposed work includes forecasting, theoretical analysis, and extending or critiquing existing approaches.
III. CONCLUSION
The paper argues that AI’s growing capabilities create both unprecedented potential benefits and greater societal impact. It presents robust and beneficial AI research as a way to maximize those benefits while avoiding pitfalls.
- Conclusion: AI researchers should ensure that AI’s future impact is beneficial while investigating how to maximize benefits and avoid potential pitfalls.The authors reject characterizing this research agenda as anti-AI and express confidence that beneficial outcomes are possible.
IV. AUTHORS
The authors bring backgrounds spanning computer science, future-of-humanity research, and physics-informed work on artificial intelligence. Their profiles connect the paper to both technical AI research and broader questions about AI’s future.
- Authors: Stuart Russell is a UC Berkeley computer science professor whose research covers artificial intelligence and machine learning.He coauthored Artificial Intelligence: A Modern Approach, described as a standard text.
- Authors: Daniel Dewey is an Oxford Future of Humanity Institute research fellow focused on machine superintelligence and AI’s future.His prior experience includes Google, Intel Labs Pittsburgh, and Carnegie Mellon University.
- Authors: Max Tegmark is an MIT physics professor studying connections between information processing in biological and engineered systems through physics-based techniques.He is also president of the Future of Life Institute, which supports robust and beneficial AI research.