Source-linked AI summary
The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, Peter Eckersley, Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff, Bobby Filar, Hyrum Anderson, Heather Roff, Gregory C. Allen, Jacob Steinhardt, Carrick Flynn, Seán Ó hÉigeartaigh, SJ Beard, Haydn Belfield, Sebastian Farquhar, Clare Lyle, Rebecca Crootof, Owain Evans, Michael Page, Joanna Bryson, Roman Yampolskiy, Dario Amodei
TL;DR
Malicious uses of AI create security threats across multiple domains, so this report surveys those threats, analyzes their evolution, and proposes prevention and mitigation strategies. It finds that AI can augment both attacks and defenses while changing the attack surface, although the long-term attacker–defender equilibrium remains unresolved.
Problem
The report addresses limited understanding of potential security threats arising from malicious uses of artificial intelligence.
Method
The report surveys threats and analyzes how AI may alter digital, physical, and political security domains.
Results
AI can augment attacks and defenses in cyberspace while changing the attack surface that hackers can target.
Takeaways & Limitations
Policymakers and technical researchers should collaborate on investigating, preventing, and mitigating potential threats while treating AI as dual-use.
Takeaways & Limitations
The report analyzes but does not conclusively resolve the long-term equilibrium between attackers and defenders.
Abstract
from arXiv · showhide
This report surveys the landscape of potential security threats from malicious uses of AI, and proposes ways to better forecast, prevent, and mitigate these threats. After analyzing the ways in which AI may influence the threat landscape in the digital, physical, and political domains, we make four high-level recommendations for AI researchers and other stakeholders. We also suggest several promising areas for further research that could expand the portfolio of defenses, or make attacks less effective or harder to execute. Finally, we discuss, but do not conclusively resolve, the long-term equilibrium of attackers and defenders.
Executive … carrying out cyberattacks will alleviate the existing tradeoff
The report examines malicious uses of AI across security domains, proposes ways to forecast, prevent, and mitigate resulting threats, and offers four high-level recommendations. It focuses on near-term attacks while leaving the long-term attacker–defender equilibrium unresolved.
- Executive: AI capabilities are growing rapidly and have widespread beneficial applications, but malicious uses have received comparatively less attention.Examples of beneficial applications include machine translation and medical image analysis.
- Summary: The report surveys potential security threats from malicious AI uses and proposes ways to forecast, prevent, and mitigate them.It focuses on attacks likely to emerge soon if adequate defenses are not developed.
- malicious uses of AI.: The report does not conclusively resolve the long-term equilibrium between attackers and defenders.Instead, it emphasizes near-term attack types and defensive needs.
- harmful applications are foreseeable.: The authors make four recommendations: policymakers and technical researchers should collaborate, AI researchers should address dual-use risks, mature practices should be adapted, and stakeholder participation should expand.The recommendations cover investigation, prevention, mitigation, research norms, best-practice transfer, and broader domain expertise.
- of AI.: As AI becomes more powerful and widespread, attacks may become cheaper and scalable, expanding existing threats and the set of capable actors and targets.AI systems may perform tasks that ordinarily require human labor, intelligence, and expertise.
- use of AI systems to complete tasks that would be otherwise: AI may introduce new attacks by enabling tasks that are impractical for humans, while malicious actors may also exploit vulnerabilities in defensive AI systems.The report distinguishes these novel threats from the expansion of existing ones.
- and likely to exploit vulnerabilities in AI systems.: The analysis separately considers three security domains and illustrates potential threat changes through representative examples.Digital security is identified as one domain, including automated tasks involved in attacks such as spear phishing.
exploit human vulnerabilities (e.g. through the use of speech … regulatory responses.
The report identifies expanding AI-enabled threats across digital, physical, and political security, including attacks on AI systems, autonomous physical operations, surveillance, persuasion, deception, privacy, and social manipulation. It recommends coordinated technical, institutional, cultural, and policy responses, supported by research into cybersecurity, openness, responsibility, and technological and regulatory interventions.
- (e.g. through automated hacking), or the vulnerabilities: AI may expand digital threats through automated hacking and novel attacks that subvert cyber systems, including attacks on AI systems.The passages also reference adversarial examples and data poisoning as relevant attack mechanisms.
- carrying out attacks with drones and other physical systems: AI may enable physical attacks by automating drone swarms, autonomous weapons, and other systems that are infeasible to direct remotely.The cited passages describe remotely operating thousands of micro-drones and using AI to automate physical-security tasks.
- surveillance (e.g. analysing mass-collected data), persuasion: Political-security threats include mass surveillance, targeted persuasion, manipulated videos, deception, privacy invasion, and social manipulation.The passages specifically connect these capabilities with analysing mass-collected data, creating targeted propaganda, and manipulating videos.
- privacy invasion and social manipulation. We also expect novel: These political uses may influence human behaviors, moods, and beliefs, with especially significant concerns for authoritarian states and possible effects on truthful democratic debate.The supplied passages state that such developments may undermine democracies’ ability to sustain truthful public debates.
- intersection of cybersecurity and AI attacks: The report proposes four research priorities: learning from cybersecurity, exploring openness models, promoting responsibility, and developing technological and policy solutions.The openness agenda includes pre-publication risk assessment, safety-favoring sharing regimes, central access licensing, and lessons from other dual-use technologies.
- technical areas of special concern, central access licensing: The proposed openness and safety measures include formal verification, responsible vulnerability disclosure, pre-publication assessment, safety-oriented sharing, and central access licensing.The report frames these measures as part of reimagining norms and institutions around the openness of dual-use AI and machine learning.
- well as policy interventions, that could help build a safer future; regulatory responses.: A safer future also requires education, ethical standards and norms, monitoring of AI-relevant resources, coordinated public-good security, and action by researchers, companies, legislators, civil servants, regulators, security researchers, and educators.The report emphasizes that the challenge is daunting and that the stakes are high.
Introduction … AI Capabilities
The report examines how rapidly advancing AI and machine learning may enable malicious attacks against digital, physical, and political security. It defines its scope around intentional malicious use of currently available or near-term AI, surveys related literature, and highlights both major capabilities and limits relevant to forecasting threats.
- Introduction: AI systems perform tasks commonly thought to require intelligence, while machine learning systems improve task performance through experience.
- Introduction: AI and machine learning have progressed rapidly, enabling applications including speech recognition, machine translation, spam filtering, search engines, and emerging autonomous systems.
- Introduction: Malicious use encompasses practices intended to compromise the security of individuals, groups, or society, including threats to digital, physical, and political security.
- Scope: The report focuses on intentional deployment or compromise of currently available or plausible-within-5-years AI systems, especially those leveraging machine learning.
- Scope: The scope excludes indirect societal effects and system-level threats arising from interactions among non-malicious actors, including AI-safety races and escalating autonomous weapons conflicts.
- Related Literature: Although related work examines specific AI-security risks and adjacent fields, the intersection of AI and malicious intent had not been analyzed comprehensively.
- AI Capabilities: Recent gains reflect more computing power, improved algorithms, software frameworks, larger datasets, and increased commercial investment, but rapid progress is concentrated in especially tractable tasks.
- AI Capabilities: Image-recognition accuracy improved from around 70% to 98%, exceeding the human benchmark of 95%, while AI-generated images became nearly indistinguishable from photographs.
NEC UIUC … Expanding Existing Threats
AI is dual-use, increasingly efficient, scalable, capable, anonymous, and rapidly diffused, while retaining distinctive vulnerabilities. Absent adequate defenses, progress may expand existing threats by broadening participation, increasing attack rates and targets, and increasing willingness to attack.
- NEC UIUC: AI systems have steadily expanded their capabilities and can exceed human performance on some tasks.The paper notes that AI performs well on a growing portion of human-capable tasks and can surpass even the most talented humans after reaching human-level performance.
- Security-Relevant Properties of AI: AI is dual-use: systems and design knowledge can support both beneficial and harmful ends.Because some intelligence-requiring tasks are benign while others are not, researchers cannot always avoid producing systems that may be directed toward harm.
- Security-Relevant Properties of AI: Once trained, AI can perform tasks more cheaply or quickly than humans, scale through computing power or copies, and diffuse rapidly through software and open research.Relevant algorithms are often reproduced within days or weeks, and papers frequently include source code.
- Security-Relevant Properties of AI: AI also introduces unresolved vulnerabilities, including data poisoning, adversarial examples, and flaws in autonomous-system goals.These vulnerabilities differ from traditional software flaws and can cause failures unlike those a human would make.
- General Implications for the Threat Landscape: Absent adequate defenses, progress in AI will expand existing threats and alter their typical character.The report expects attacks to become more effective, finely targeted, harder to attribute, and more likely to exploit AI-system vulnerabilities.
- Expanding Existing Threats: AI may expand familiar attacks by increasing the number of capable actors, attack rates, and plausible targets.This follows from AI’s efficiency, scalability, and ease of diffusion, which can reduce the resources needed to conduct attacks.
- Expanding Existing Threats: Automating spear-phishing research and message generation could enable more actors, cross-language operations, and mass spear phishing.Personalized attacks currently require skilled target research and contextual message creation, but automation could make them less discriminate in target selection.
- Expanding Existing Threats: Greater anonymity and psychological distance may increase actors’ willingness to carry out attacks.Perceived untraceability and reduced empathy or anticipated trauma can lower barriers to attacking others; robotics and cheaper hardware also contribute to threat expansion.
Introducing New Threats … Digital Security
AI progress may enable attacks that exceed human capabilities or exploit vulnerabilities specific to AI systems. In digital security, these changes could make attacks more effective, targeted, anonymous, adaptive, and scalable across social engineering, vulnerability discovery, hacking, and other cyber-offenses.
- Introducing New Threats: AI could enable attacks that are otherwise infeasible by exceeding human capabilities, including realistic voice imitation and autonomous control of robot swarms or malware.AI systems may perform tasks more successfully than any human could, or control systems that humans cannot manually direct.
- Introducing New Threats: Novel AI systems may create attack opportunities by exposing unresolved vulnerabilities, such as adversarial examples that cause self-driving cars to crash.These vulnerabilities can be exploited because AI systems may interpret altered inputs differently from humans.
- Altering the Typical Character of Threats: AI-supported attacks are expected to become especially effective, finely targeted, difficult to attribute, and exploitative of vulnerabilities in AI systems.The threat landscape may expand through both stronger existing threats and new threats that do not yet exist.
- Altering the Typical Character of Threats: Efficiency, scalability, and AI capabilities may make highly effective and finely targeted attacks more typical, while increasing anonymity may make attacks harder to attribute.AI can help attackers identify and analyze targets, tailor attacks to target properties, and use autonomous systems instead of acting in person.
- Scenarios: The scenarios illustrate plausible malicious uses of AI across digital, physical, and political security rather than definitive forecasts.Some scenarios may not be technically possible within 5 years or may not be realized even if technically possible.
- Digital Security: In digital security, AI could automate personalized social engineering by generating tailored malicious content, impersonating contacts, and sustaining trust-building dialogues.Attackers could use victims’ online information to create likely-to-be-clicked websites, emails, and links, potentially including visual impersonation in video chats.
- Digital Security: AI-enabled digital attacks could automate vulnerability discovery, target prioritization, exploit generation, evasion, adaptive hacking, denial-of-service, criminal service tasks, data poisoning, and model extraction.The examples include estimating victims’ wealth and willingness to pay, poisoning consumer models, inferring remote model parameters, and using human-like autonomous agents to overwhelm services.
- Digital Security: Neural-network-augmented fuzzing may strengthen well-defended systems before organized crime adopts the techniques to produce continuously updated ransomware exploits against enduringly vulnerable devices.Fully patched operating systems and browsers are mostly resistant, while many older phones, laptops, and IoT devices remain vulnerable.
Physical Security
AI can repurpose commercial autonomous systems for physical attacks, lowering the expertise needed for some attacks while increasing their scale, coordination, and distance from perpetrators.
- Commercial AI systems can be repurposed for harmful physical uses, including drones or autonomous vehicles delivering explosives and causing crashes.
- An intruding cleaning robot used visual detection to approach a finance minister and trigger a concealed explosive, killing the minister and wounding nearby staff.
- AI-enabled automation can endow low-skill individuals with previously high-skill attack capabilities by reducing expertise requirements for attacks such as self-aiming, long-range sniper rifles.
- Human-machine teaming with autonomous systems can increase attack scale, such as one person launching many weaponized autonomous drones.
- Cooperating autonomous robotic networks can provide broad surveillance and execute rapid, coordinated attacks at machine speed.
- Autonomous operation can further separate physical attacks from their initiators in time and space, including when remote communication is impossible.
Political Security
AI can expand state surveillance to massive scales, including suppression of debate, while predictive systems may flag individuals as potential threats. Political misuse also includes fabricated media, targeted disinformation, influence campaigns, information flooding, and algorithmic manipulation of content access.
- State surveillance: Automated image and audio processing can scale state surveillance for intelligence collection and the suppression of debate.The report describes automation as extending nations’ surveillance powers across collection, processing, and exploitation of intelligence information.
- Predictive policing: Predictive civil disruption systems may flag individuals as potential threats before disruptive activity occurs.The example presents police acting on such a flag and claiming 99.9% accuracy.
- Information manipulation: Political manipulation can use fabricated leader videos, hyper-personalised voter messages, influencer targeting, bot-generated noise, and content-curation algorithms.These mechanisms can affect voting behavior, exploit social-network analysis, swamp information channels, or steer users toward or away from selected content.
Security … How AI Changes The Digital Security Threat Landscape
The paper examines how AI may alter digital security by expanding the scale, number, diversity, adaptability, and evasiveness of cyberattacks, while also improving defensive capabilities. It argues that technical and policy innovations are needed to ensure AI’s net impact on digital systems is beneficial.
- Domains: The report distinguishes digital-security threats to the confidentiality, integrity, and availability of digital systems from physical- and political-security threats.It analyzes these domains separately, describing existing attack and defense conditions before considering how AI progress and diffusion may change attack nature or severity.
- Security: AI-enabled defenses are developing alongside offensive applications, but further technical and policy innovations are needed for AI’s impact on digital systems to be net beneficial.The report frames this need as preparation against increases in attack number, scale, and diversity resulting from contemporary and near-term AI.
- Context: Cybersecurity is labor-constrained and already uses AI defensively for anomaly, spam, and malware detection, creating strong opportunities for automation.82% of surveyed decision-makers reported a shortage of needed cybersecurity skills, while malicious actors have incentives based on speed, labor costs, and difficulty retaining skilled workers.
- Context: Publicly disclosed offensive AI use has been limited to white-hat experiments, but machine-learning-enabled cyberattacks may soon appear in the wild or already be used by sophisticated adversaries.The report notes that circumstantial evidence has supported claims that motivated adversaries may already be using AI offensively.
- How AI Changes The Digital Security Threat Landscape: AI may let attackers conduct larger-scale, more numerous, and more diverse attacks with the same skill and resources, as demonstrated by automated spear phishing that tailored malicious tweets to users’ interests.The system achieved a high click-through rate to a potentially malicious link.
- How AI Changes The Digital Security Threat Landscape: AI adaptability could change the offense-defense balance because machine-learning systems have learned to evade endpoint detection and response platforms that are effective against typical human-authored malware.EDR platforms combine heuristic and machine-learning algorithms for behavioral analytics, next-generation antivirus, and exploit prevention.
- How AI Changes The Digital Security Threat Landscape: Researchers have used machine learning to generate command-and-control domains indistinguishable from legitimate domains and reinforcement learning to manipulate malicious binaries for evasion.These techniques aim to help malware communicate with host machines while avoiding detection.
- How AI Changes The Digital Security Threat Landscape: Attackers are expected to use reinforcement learning to craft experience-driven attacks that current technical systems and IT professionals are unprepared for without additional investment.The report expects cybercriminals and state actors to adopt these techniques as capable AI becomes more widely distributed, including unexplored offensive applications.
Points of Control and Existing Countermeasures
Cybersecurity can be mitigated through multiple control points, but existing countermeasures have important limitations and require proactive adaptation as AI and cybersecurity evolve together. These limitations include difficult attribution and enforcement, risks from centralized and AI-based defenses, and unequal capacity to defend against physical attacks.
- Points of Control and Existing Countermeasures: Multiple control points can increase security, but proactive efforts are needed because AI and cybersecurity will rapidly evolve in tandem.The report identifies existing countermeasures while noting that potential interventions remain unproven.
- Consumer awareness: Consumer awareness improves phishing detection and security habits, yet most users remain vulnerable to simple attacks exploiting unpatched systems.Examples of better practices include diverse, complex passwords and two-factor authentication.
- Industry centralization: Centralized systems such as spam filters and network monitoring benefit from economies of scale, but compromise raises stakes and attackers can study defenses to evade them.Attackers may analyze commercial antivirus updates to identify what protections do and do not cover.
- Governments and researchers: Laws, norms, and responsible disclosure can support defense, but researcher risks, cross-border enforcement, attribution difficulties, and conflicting incentives limit their effectiveness.Attribution may be necessary for deterrence and punishment, while revealing evidence can compromise intelligence sources or methods.
- Technical cybersecurity defenses: Technical defenses include patching, threat detection, incident response, vulnerability discovery, and machine-learning detection, but their relative effectiveness remains poorly analyzed.Machine-learning defenses use supervised learning to generalize from known threats or unsupervised learning to detect anomalous behavior.
- Technical cybersecurity defenses: AI-based cyber defenses can expand the attack surface because they are insufficiently hardened against anticipating attackers, while professionals report low confidence in them.The report therefore calls for further development of these technologies.
How AI Changes the Physical Security Landscape
AI can magnify physical-security threats by making customizable robots autonomous, enabling more precise, persistent, large-scale, and remotely conducted attacks. Autonomous robots and cyber-physical systems also introduce vulnerabilities that are difficult to fully prevent because defenses are costly and imperfect.
- Threats from AI-enabled robots: Robots equipped with dangerous payloads can conduct precise physical attacks from long distances, while autonomy magnifies this threat beyond existing human-piloted systems.This capability was previously limited to actors able to afford technologies such as cruise missiles.
- Threats from AI-enabled robots: Autonomy can let one person cause greater damage with robots, enabling large-scale attacks by smaller groups using increasingly mature open-source software.Relevant tools include face detection, navigation and planning algorithms, and multi-agent swarming frameworks.
- Threats from AI-enabled robots: Long operating durations, specialized sensing, and tolerance of smoke, toxins, darkness, fog, or underwater conditions allow robots to sustain attacks and reach environments humans cannot.These capabilities can enable attacks or keep targets at risk over extended periods.
- Cyber-physical vulnerabilities: Remote manipulation, traditional cybersecurity flaws, adversarial examples, and misplaced trust create acute risks for autonomous robots and high-risk cyber-physical systems.Examples include hacked service robots, insecure IoT systems, self-driving cars, and autonomous weapons.
- Points of Control and Existing Countermeasures: Physical defenses and governance can reduce harm, but capital-intensive, imperfect defenses and resistance to strong bans may leave an extended period of difficult-to-prevent AI-enabled physical attacks.The report identifies hardware manufacturers, distributors, software suppliers, and users as potential control points.
How AI Changes the Political Security Landscape · Interventions
AI may reshape political security through scalable manipulation, surveillance, content control, and attacks that undermine institutions, while defensive measures remain incomplete. The report recommends closer technical-policy collaboration, dual-use responsibility, adapted security practices, broader stakeholder engagement, and further investigation.
- How AI Changes the Political Security Landscape: AI could intensify political manipulation through increasingly sophisticated autonomous actors, misleading information, targeted messaging, and synthetic multimedia.Existing AI techniques can generate convincing text, while systems may target individuals at moments of maximum persuasive potential.
- How AI Changes the Political Security Landscape: Reduced trust in online information may advantage groups that thrive in low-trust societies, particularly authoritarian regimes that devalue objective truth.Authoritarian governments may also use AI for fine-grained surveillance and network analysis to prioritize attention toward potential subversive leaders.
- How AI Changes the Political Security Landscape: AI-enabled filtering and prior digital or physical attacks could manipulate public opinion, disrupt political institutions, or justify more authoritarian policies.Technical tools may reside with states in authoritarian regimes and corporations in democracies.
- How AI Changes the Political Security Landscape: Existing countermeasures include bot and forgery detection, certified media authenticity, encryption, campaign-targeting transparency, and discourse-improvement proposals, but none definitively solves the problems.Detecting misleading news and images remains unsolved while generation of apparently authentic multimedia and text advances rapidly.
- Interventions: The report presents four high-level recommendations focused on strengthening dialogue among technical researchers, policymakers, and other stakeholders.It emphasizes exploratory investigation rather than highly specific technical or policy proposals.
- Interventions: Policymakers should collaborate with technical researchers, while AI researchers should account for dual-use risks and proactively address foreseeable harmful applications.Policy should avoid impeding research unless restrictions provide commensurate benefits.
- Interventions: The broader AI community should adapt mature dual-use practices such as red teaming and expand discussions to civil society, security experts, businesses, ethicists, and the public.The report also anticipates adaptive defensive action by citizens, whose ability to respond may vary with technological literacy.
Priority Areas for Further Research … Developing Technological and Policy Solutions
The report proposes a research agenda spanning cybersecurity practices, openness models, institutional responsibility, and technological and policy defenses against malicious AI use. It emphasizes continued investigation, transparent and proportionate governance, and sustained support from qualified organizations and funders.
- Learning from and with the Cybersecurity Community: Cybersecurity should remain a major, ongoing priority for preventing and mitigating AI-related harms, with applicable security best practices transferred to AI systems.The report links increasing AI-enabled control of physical systems to growing cybersecurity risks.
- Learning from and with the Cybersecurity Community: Research priorities include red teaming, formal verification, responsible disclosure of AI vulnerabilities, security tools, and secure hardware.The report highlights extensive red teaming for critical systems and asks how verification, vulnerability disclosure, testing tools, and hardware security features should be developed.
- Exploring Different Openness Models: Because openness enables collaboration and progress but can increase malicious actors’ access to powerful capabilities, researchers should investigate when publication should be limited or delayed.The report does not propose a specific solution and stresses that reduced openness must remain sensitive to openness’s recognized benefits.
- Exploring Different Openness Models: Any reconsideration of openness should be transparent, publicly justified, and no broader than necessary to address genuine misuse risks rather than corporate competitiveness.Potential mechanisms include pre-publication risk assessment for technically sensitive areas such as digital security and adversarial machine learning.
- Promoting a Culture of Responsibility: AI researchers and their organizations should deepen a culture of responsibility by learning from other technical fields and focusing more closely on malicious-use risks.Suggested initiatives include education on ethical and socially responsible technology use and ethical statements and standards.
- Developing Technological and Policy Solutions: Further research should examine privacy protection, coordinated AI use for public-good security, monitoring, and other institutional and technological defenses against misuse.The report asks how technical and institutional measures can protect privacy, distribute defensive AI, and support security best practices.
- Developing Technological and Policy Solutions: Implementing the research agenda requires incentives, relevant expertise, organizational commitment, proven track records, and additional public and private funding.The report presents raising awareness and laying out an initial agenda as an initial step, while identifying further expertise and monetary resources as necessary.
Analysis … Conclusion
AI will become increasingly central to security, with malicious uses expanding across digital, physical, and political domains while scalable defenses, regulation, and coordination shape uncertain outcomes. Preparing for this interconnected and rapidly changing security landscape is urgent despite unresolved disagreements and long-term uncertainty.
- Analysis: Long-term security outcomes remain uncertain because technologies, malicious strategies, stakeholder responses, and policy factors may evolve beyond current forecasts.Even a seemingly stable medium-term equilibrium could be short-lived, and developments unrelated to AI may prove more impactful.
- Attacker Access to Capabilities: Open access to cutting-edge research is expected to significantly increase attackers’ ability to cause harm with digital and robotic systems over the next five years.This expectation follows from AI’s dual-use nature, efficiency, scalability, and ease of diffusion.
- Attacker Access to Capabilities: Developers and regulators are expected to impose more limits on access to or malicious use of powerful AI, but the effectiveness of restricting or monitoring access remains uncertain.The report also identifies preemptive design and novel organizational and technological measures within international policing as relevant responses.
- Existence of AI-Enabled Defenses: AI’s characteristics can support scalable defenses, including improved malware detection, counter-drone systems, and expanded use of AI in criminal investigations and counterterrorism.Some defenses may nevertheless be prohibitively expensive, making the pace of technical development and deployment important.
- Distribution and Generality of Defenses: The largest potential harms may be least likely when governments and major corporations can adopt coordinated defenses, while other victims require defenses embedded in widespread or low-cost technologies.Reliance on fortified platforms may increase concentration of data and power, and tailored defenses may require substantial financial backing.
- Distribution and Generality of Defenses: Misaligned incentives may leave available defenses unused, requiring regulation or other approaches to protect individuals who cannot directly improve security.Examples include data-breach victims and people affected by DDoS attacks using botnets.
- Overall Assessment: Malicious use of AI is expected to increase with society’s broader adoption of AI, while defense prospects include securing AI systems and improving vulnerability discovery, but manipulation attacks may remain especially difficult to address.Attribution and attacker penalization may also remain difficult without significant effort, and technology and media giants may become privileged providers of tailored protection.
- Conclusion: AI will figure prominently in future security, connect digital, physical, and political security more closely, and create urgent preparation needs as social engineering and other attacks become more sophisticated.The report urges precautionary action and collective efforts to understand commonalities across malicious uses and improve prevention and mitigation.
Appendix A: … Red Teaming
The appendix documents a multidisciplinary workshop and report-writing process, then outlines research directions on dual-use governance and red teaming to improve understanding and mitigation of malicious AI uses.
- Event Structure: The workshop combined expert presentations, breakout discussions of security scenarios and defenses, and prioritization of useful and tractable prevention and mitigation measures.Participants agreed that a research agenda should guide subsequent report writing.
- Report Writing Process: The report drew on workshop notes and the authors’ prior and subsequent research, with drafts circulated to attendees and additional domain experts.The authors noted that the final report could not capture every participant perspective.
- Research: The appendix supplements the main report’s recommendations and priority research areas with initial questions, directions, and threat-factor links for researchers.It is presented as a jumping-off point rather than a definitive research program.
- Dual Use Analogies and Case Studies: Dual-use case studies suggest combining soft norms and hard laws while learning from both successful governance and failures such as cryptography export controls.AI’s similarities to cryptography may make comparable controls tempting, but the report advises approaching that path cautiously.
- Dual Use Analogies and Case Studies: Open dual-use questions concern the appropriate governance level, transferable norms, AI-specific challenges, exemplary cases, and lessons from failed controls.The questions span fields, algorithms, hardware, software, and data.
- Red Teaming: Red teaming uses controlled attacks by security experts or organizational members, optionally opposed by a blue team, to reveal how systems and practices can be improved.AI-enabled cyber offense and defense and adversarial machine learning appear especially suitable, although broader AI red teaming is also beneficial.
- Red Teaming: Red teaming could test hypothetical AI cyberattacks and real-world machine-learning vulnerabilities, while raising questions about detection coverage, responsibility, skills, incentives, lesson uptake, and extension to physical and political domains.Examples include broader cyberattack ranges, CleverHans benchmarks, and the NIPS 2017 adversarial attacks and defenses competition.
Formal Verification … Security Tools
Formal methods may improve the security of AI systems, but end-to-end verification is constrained by complexity, specification difficulties, and limited real-world adoption. Complementary measures include responsible vulnerability disclosure, exploit bounties, and development tools targeting AI-specific attacks.
- Formal Verification: Formal verification of AI systems may be difficult or impossible end to end because modern systems are complex and desired properties can be hard to specify.Verification would ideally establish that internal processes attain specified goals, goals remain constant against adversarial changes, and deception by adversarial inputs is bounded.
- Verifying Hardware: Component-level formal methods may remain feasible even when complete AI-system verification is prohibitively expensive, with hardware especially amenable to verification.Formal methods have been widely adopted in the hardware industry for decades.
- Verifying Security: Formal verification can provide robust safety guarantees for some security protocols, but tools such as CryptoVerif remain largely theoretical and have limited real-world adoption.CryptoVerif lets programmers check code correctness during development.
- Verifying AI Functionality: Verification of selected AI properties, including image classifiers and adversarial examples in regions, remains feasible despite state-space explosion in arbitrary complex systems.The feasibility of verifying components does not imply that whole-system behavior is tractable.
- Responsible “AI 0-Day” Disclosure: AI vulnerabilities include adversarial misclassification, training-data poisoning, and traditional flaws such as memory overflow, motivating research into responsible disclosure norms.Open questions concern whether cybersecurity’s disclosure norm should extend to AI and whether AI systems should be presumed vulnerable until proven secure.
- Responsible “AI 0-Day” Disclosure: Responsible disclosure research must address notification recipients, publication timing, institutional processing, AI-equivalent patching, and trade-offs among resource demands, accuracy, and robustness to noise.These choices must account for rapidly changing attacks and defenses.
- AI-Specific Exploit Bounties: AI-specific exploit bounties could extend existing vendor programs or be offered by governments, NGOs, or philanthropies when vendors are unwilling or unable to provide them.The question is especially relevant to popular open-source machine-learning frameworks and academic projects.
- Security Tools: Security tools for AI development and deployment could reduce attackability through adversarial-data generation, classification-error analysis, extraction detection, vulnerability scanning, and robustness suggestions.These tools would extend existing software capabilities such as testing, fuzzing, and anomaly detection.
Secure Hardware · Pre-Publication Risk Assessment in Technical Areas of Special Concern · Central Access Licensing Models
The report proposes secure AI hardware, pre-publication risk assessment, and central access licensing as approaches to reduce malicious use while preserving some beneficial access and openness. Each approach has important limitations and open questions, including cost, extraction, concentration, feasibility, and trade-offs.
- Secure Hardware: Secure AI hardware could restrict copying, control the number of AI systems, and support hardware-level access restrictions and audits.A proposed feature would prevent copying a trained model without first deleting the original copy.
- Secure Hardware: Tamper-proof hardware could deter duplication, reveal tampering, and hard-code safe operating properties, but secure processors cost significantly more and lack AI-specific development.These possibilities are described as potentially valuable for protecting AI systems and enabling credible commitments to safe operation.
- Secure Hardware: Open questions for secure hardware concern AI-specific requirements, adoption incentives, compliance across international vendors, applicability of existing designs, secure enclaves, and affordability.The report also asks whether changing AI risks justify a major hardware-security overhaul and how policy mechanisms might encourage use despite cost premiums.
- Pre-Publication Risk Assessment in Technical Areas of Special Concern: Pre-publication risk assessment evaluates the risks of widely available capabilities and uses that analysis to decide whether, and how extensively, to publish them.The report notes that this norm is already widespread in computer security, where proofs of concept may be published instead of fully working exploits.
- Pre-Publication Risk Assessment in Technical Areas of Special Concern: Publication openness can vary from rough ideas to source code, trained models, and practical guidance, while voice synthesis illustrates a capability whose full publication may be considered too risky.The report identifies potential criminal applications including automated spearphishing and disinformation.
- Pre-Publication Risk Assessment in Technical Areas of Special Concern: Because machine learning benefits from prevalent openness, restrictions should be weighed carefully, especially if restricted publication becomes common rather than rare.The report compares this concern with restricted publication in biotechnology and vulnerability disclosure in cybersecurity research.
- Central Access Licensing Models: Central access licensing offers capabilities through secure, remotely accessible infrastructure while withholding underlying code and applying terms and conditions.It can provide universal access while keeping technological breakthroughs away from bad actors, although this also excludes well-intentioned researchers.
- Central Access Licensing Models: Centralized access may constrain large-scale attacks and query-based attacks, but black-box extraction and concentration of services can create security risks and insider threats.The report highlights model extraction, concentrated organizational power, and questions about query limits, provider vetting, malicious-use detection, and technological, legal, and political feasibility.
Sharing Regimes that Favor Safety and Security … Differential privacy
The paper proposes controlled information sharing, education, ethical standards, beneficial-AI norms, and privacy technologies as complementary ways to reduce security risks while preserving useful AI development. It also identifies open implementation questions and trade-offs, including the performance costs of differential privacy.
- Sharing Regimes that Favor Safety and Security: Trusted parties could selectively share AI capability information and data to support collaborative safety and security analysis while limiting risks from broader diffusion.The approach is analogous to cyber-domain ISACs and ISAOs, with antivirus and large technology companies serving as knowledge-sharing concentrations.
- Sharing Regimes that Favor Safety and Security: Capabilities could be published after analysis concludes that further diffusion would not cause harm, but the model raises questions about trust, information scope, incentives, and risks.The proposals also overlap partially with pre-publication risk assessment and central access licensing.
- Security, Ethics, and Social Impact Education for Future Developers: Ethics education could help AI researchers recognize malicious-application risks and make more informed decisions about technology openness and design.The paper notes that long-term research on education’s effects on researchers’ careers and eventual decisions is still lacking.
- Ethics Statements and Standards: Companies, research organizations, and other deployers could develop and sign multi-stakeholder ethical standards, building on initiatives such as IEEE and the Asilomar AI Principles.Open questions concern implementation, accountability, organizational statements, industry-wide dialogue, and standards that remain updateable yet objective.
- Norms, Framings and Social Incentives: Normative framing could emphasize AI’s mutually beneficial security upsides and discourage harmful exploitation despite risks that may benefit individual actors.The paper asks what historical analogies and governance processes could support a beneficial normative culture for AI.
- Technologically Guaranteed Privacy: Algorithmic privacy technologies could protect private information relevant to threats such as automated spear phishing and personalized propaganda, alongside procedural and legal safeguards.The paper highlights differential privacy as a potentially relevant technology and asks how to encourage its adoption in AI systems.
- Differential privacy: Publicly deployed machine-learning models can expose information from underlying datasets through model inversion or membership inference attacks, even without training-data access.These attacks may allow individuals to break dataset anonymity by querying the model in particular ways.
- Differential privacy: Differentially private algorithms add noise to training data to provide strong information-leakage guarantees while minimizing performance effects, but they generally lose performance relative to non-private models.This trade-off may undermine privacy if model developers lack incentives to keep datasets private.
Secure Multi-Party Computation · Monitoring Resources · Exploring Legal and Regulatory Interventions
The section presents secure multi-party computation as a way to preserve privacy while enabling collaborative machine learning, web applications, cloud computation, and surveillance, though its computational overhead limits applicability. It also considers monitoring AI inputs and legal or regulatory interventions, emphasizing unresolved design questions and the need for coordination across technical, policy, and legal communities.
- Secure Multi-Party Computation: Secure multi-party computation lets parties jointly compute functions without revealing their private inputs.A simple example is jointly computing a vote without sharing individual votes.
- Secure Multi-Party Computation: MPC can train machine-learning systems on sensitive data without significantly compromising privacy.Examples include researchers training on confidential patient records through a hospital and companies learning from user data without accessing it directly.
- Secure Multi-Party Computation: MPC supports privacy-preserving web applications, cloud computation, and surveillance by avoiding collection or access to sensitive data.Individuals can receive predictions based on health data without sending copies to the company, while surveillance systems can classify faces or web activity without accessing the underlying data.
- Secure Multi-Party Computation: MPC’s computational overhead can increase by multiple orders of magnitude, making it best suited to relatively simple computations or especially privacy-sensitive uses.This limitation constrains broader deployment despite MPC’s privacy benefits.
- Monitoring Resources: Monitoring inputs to AI systems, especially computing hardware, could help predict or prevent misuse, drawing on monitoring regimes for fissile materials and chemical production.Open questions concern the feasibility, tractability, importance, consequences, drawbacks, and choice of AI inputs to monitor.
- Exploring Legal and Regulatory Interventions: Legal and regulatory interventions should be considered carefully because ill-considered government action could be counterproductive.The paper raises questions about responsibility, institutional roles, liability, coordination, preparedness, and policies governing privacy-preserving defenses and malicious AI use.
- Exploring Legal and Regulatory Interventions: Long-term effectiveness requires trust and proactive coordination between the AI community and legal and policy institutions.The paper argues that reactive responses by separate sectors are likely to produce clumsy, ineffective outcomes, motivating Recommendations #1 and #2.