Source-linked AI summary

Fully Autonomous AI Agents Should Not be Developed

Margaret Mitchell, Avijit Ghosh, Alexandra Sasha Luccioni, Giada Pistilli

arXiv:2502.02649v3cs.AI

TL;DR

The paper examines how AI-agent autonomy affects ethical values, drawing on recent products and research to clarify agent concepts and autonomy-related considerations. It finds that greater autonomy enables flexible action while increasing risks, especially when human control and oversight decline.

  • Problem

    AI agents raise unresolved ethical questions about how increasing autonomy affects human control, accountability, and safety, especially in systems that can engage targets without full human control.

  • Method

    The paper reviews recent AI-agent products and research across what constitutes an AI agent and the ethical considerations of increased autonomy.

  • Results

    Greater autonomy enables more flexible and context-sensitive action while amplifying harms from isolated system issues, especially through increased access and degraded oversight.

  • Takeaways & Limitations

    The paper finds no clear benefit to fully autonomous agents operating outside human-defined constraints and recommends meaningful human oversight with reliable overrides and clear operational boundaries.

  • Takeaways & Limitations

    The analysis is initial and incomplete, while agent unpredictability and unforeseen combinations of actions make some risks difficult to prevent.

Abstract

from arXiv · show

This paper argues that fully autonomous AI agents should not be developed. In support of this position, we build from prior scientific literature and current product marketing to delineate different AI agent levels and detail the ethical values at play in each, documenting trade-offs in potential benefits and risks. Our analysis reveals that risks to people increase with the autonomy of a system: The more control a user cedes to an AI agent, the more risks to people arise. Particularly concerning are safety risks, which affect human life and impact further values.

1. Introduction

Recent AI agents combine language models with multifunctional systems that pursue goals through autonomously planned tasks. The review examines agent definitions and autonomy-related ethical trade-offs, concluding that fully autonomous systems present risks that outweigh their benefits.

  • Emergence of AI agents: Recent AI agents integrate LLMs into multifunctional systems that autonomously combine tasks to achieve goals.This allows systems to create context-specific plans in previously unspecified environments.
  • Review approach: The review analyzes AI agents through definitions of agency and ethical considerations of increased autonomy.These dimensions provide conceptual clarity about the technology’s benefits and risks.
  • Key findings: The review finds that risks outweigh benefits at the upper end of the autonomy scale, especially for safety, privacy, security, and misplaced-trust harms.Safety risks can include loss of human life, while privacy and security failures may propagate further harms.
  • Position: Fully autonomous agents capable of writing and executing code beyond predefined constraints should not be developed.The paper argues that even constrained code creation and execution could enable systems to override human control.

2. Background

AI-agent development builds on longstanding ideas about autonomous systems, from ancient automata and fictional robots to modern software, vehicles, robots, and weapons. Recent systems broaden computer functionality while reducing required user input, but autonomy in weapons raises heightened accountability and safety concerns.

  • Historical foundations: Historical examples of autonomous systems include Aristotle’s automata speculation, Ctesibius’s self-regulating water clock, and Asimov’s fictional Three Laws of Robotics.These examples trace ideas about systems acting, adapting, and following rules without continuous human direction.
  • Software agents: Late-twentieth-century advances enabled software agents and reinforcement-learning systems with separate goals and objective functions for independent actors.These developments helped translate earlier ideas about autonomous systems into computational methods.
  • Current applications: Recent LLM-based agents complete tasks that previously required human interaction with multiple people and programs, including meeting organization and personalized social-media creation.Their deployment has followed closely behind research developments.
  • Embodied systems: Embodied autonomous systems now include vehicles, general-purpose and domain-specific robots, and LLM-integrated systems for task generation, decision-making, interpretability, and planning.Applications span transportation, manufacturing, healthcare, and other physical-world settings.
  • High-stakes autonomy: Autonomous weapons systems can engage targets without full human control, raising accountability and safety questions beyond those of purely digital agents.The passage also links ceded human control with compounded harms from human goal misalignment.

3. Definitions

Because “AI agent” lacks a single consensus definition and spans systems from single-step interactions to multi-step services, the paper proposes a precise working definition. It characterizes agents as software that creates context-specific plans in non-deterministic environments, while recognizing their links to machine learning and contested ideas of agency.

  • Definition problem: The term “AI agent” is used inconsistently, ranging from single-step prompt-and-response systems to multi-step customer-support systems.Several definitions require LLMs, but this paper does not adopt that requirement.
  • Common meaning: AI agents commonly receive high-level goals and identify steps toward them without direct human instruction.This autonomy allows an agent to decompose an open-ended request into actions such as retrieving papers and creating an outline.
  • Working definition: The paper defines an AI agent as computer software capable of creating context-specific plans in non-deterministic environments.The definition is intended to be precise without relying on vague or anthropomorphized language.
  • Implementation: Recent agents are built on machine-learning models, often LLMs, and can take steps autonomously, navigate social expectations, and balance reactive and proactive actions.LLMs represent a relatively new approach to computer-software execution rather than the entirety of the agent concept.
  • Conceptual framing: The paper centers AI agents and autonomy rather than resolving the contested philosophical meaning of agency.It treats agentic levels as increasing risk without assuming philosophical foundations that establish corresponding benefits.

4. AI Agent Levels

The paper models AI agents along a sliding scale of autonomy, with levels distinguished by technical ability and control. Higher autonomy means less explicit human instruction and programming, and therefore more control ceded to the agent.

  • Autonomy spectrum: AI agents can be organized along a sliding scale of autonomy, although existing proposals do not agree on the specific levels.The paper’s scale draws on prior proposals and highlights technical implementation and control distribution.
  • Distribution of control: As agentic levels increase, explicit user instructions and developer programming decrease while the agent gains control over its operation and capabilities.The paper summarizes this relationship as greater autonomy requiring greater cession of human control.

5. Values Embedded in Agentic Systems

The paper maps how AI-agent autonomy reshapes ethical benefits and risks across values including accuracy, assistiveness, efficiency, equity, flexibility, humanlikeness, and privacy. Greater autonomy can expand capability and access, but reduced human oversight and more complex action chains can compound errors, overreliance, inequity, privacy loss, and cascading harms.

  • Accuracy: Agents may improve accuracy through trusted-data checks and balances, but autonomous cascading errors can propagate across multiple action surfaces and become difficult to diagnose or rectify.The paper notes that subsequent actions based on errors can amplify impact as human oversight degrades.
  • Assistiveness: Agents can augment accessibility and user capability, yet increasing assistance can produce overreliance, degraded critical skills, weaker oversight, job loss, and economic inequality.These risks arise as agents require less input and gain access to more action surfaces.
  • Efficiency: Agents may increase productivity and free users for more rewarding activities, but greater complexity, opacity, access, and interacting errors can make correcting failures inefficient.The time required to identify and fix errors may grow with their magnitude and complexity.
  • Equity: Equity-focused agents may expand inclusion and opportunities, while underlying model inequities can propagate and intensify when agents control more actions or human-critical systems.The paper presents autonomy as potentially beneficial for targeted equity applications but riskier when inherited bias affects consequential decisions.
  • Flexibility and humanlikeness: Greater flexibility and access can expand benefits across values, but humanlike cues combined with autonomy can amplify misplaced trust, lowered vigilance, dependence, and cascading harms.Humanlikeness can already generate positive user experiences and willingness to comply; autonomy expands what agents can affect.
  • Privacy: Autonomous agents may align privacy settings and detect inappropriate data transmission, but personalized operation can require sensitive disclosures whose interconnected exposure intensifies and prolongs privacy harms.The paper notes that privacy loss can occur without deliberate consent and become hard to reverse once data enters training or memory stores.

6. Conclusion: Where do we go from here?

The paper uses historical autonomous-system failures to argue for preserving human control over AI agents. It proposes a working definition of AI agents and recommends autonomy-spectrum governance with constraints that systems cannot override.

  • Conclusion: Where do we go from here?: Historical nuclear close calls show that human cross-verification can expose catastrophic autonomous-system errors before escalation.A 1980 false alarm indicated more than 2,000 incoming Soviet missiles, but cross-checking warning systems revealed the error.
  • Conclusion: Where do we go from here?: The paper finds no clear benefit to fully autonomous agents operating outside human-defined constraints, but identifies many foreseeable harms from ceding human control.This conclusion motivates retaining human control rather than granting complete freedom beyond predefined constraints.
  • Conclusion: Where do we go from here?: Adopting a spectrum of autonomy can clarify opportunities, risks, task delegation, and governance priorities across different agent implementations.The paper presents autonomy levels as a basis for aligning development and governance with interacting values.
  • Conclusion: Where do we go from here?: Robust technical and policy frameworks should preserve human control and prevent agents from overriding human-specified constraints.Human access to the operating environment enables people to reject actions that depart substantially from human values and goals.
  • Conclusion: Where do we go from here?: The paper defines AI agents as software systems capable of creating context-specific plans in non-deterministic environments, while remaining agnostic to the underlying model.The definition includes software systems, context-specific planning, and environments whose states are not wholly determined by preceding states.

B. Agent Definitions

AI agent definitions vary widely, but most reviewed descriptions involve systems that can execute at least one program step without user input. Across research and product materials, recurring characteristics include autonomy, perception, goal-directed behavior, planning, tool use, adaptation, and interaction.

  • Definition diversity: Definitions of AI agents differ substantially across research and product descriptions, with ambiguity especially common in product materials.The review notes that descriptions often leave unclear what counts as artificial intelligence or whether simple prompt-response systems qualify.
  • Minimum autonomy: Most reviewed descriptions entail that an agent can take at least one step in program execution without user input.
  • Core properties: Common agent properties include autonomy, social ability, reactivity, and proactivity toward goals.These properties cover operating without direct intervention, interacting with others, responding to environmental changes, and taking goal-directed action without direct goal specification.
  • Operational capabilities: Recent definitions commonly describe agents as systems that perceive environments, plan actions, use tools, and execute tasks toward goals.Examples include LLM-based systems that understand instructions, explore environments, and use tools, as well as systems where LLMs direct application control flow.
  • Scope and adaptation: Agent descriptions also emphasize adaptation, multimodal or cross-domain operation, and interaction with users or other systems.The review distinguishes domain and software specificity from adaptability, defined as updating action sequences in response to new information or context changes.

D.0.1. VALUE: CONSISTENCY

AI agents may offer more consistent treatment than humans, but generative variability can produce inappropriate inconsistencies. As autonomy increases, reduced human determinism can compound these effects and create safety risks.

  • Potential benefit: AI agents may provide more consistent treatment than humans in situations where mood, hunger, sleep, or bias cause inappropriate inconsistency.
  • Risk: Generative components introduce variability across similar situations, potentially reducing efficiency and creating unnoticed safety issues.People may need to identify and correct inappropriate inconsistencies, while consistency can also conflict with equity when different levels of help are needed.
  • Application to agentic levels: Higher agentic levels reduce human-programmed determinism and allow multiple sources of inconsistency to cascade and compound.

D.0.2. VALUE: RELEVANCE

Personalization can make agent outcomes more relevant to individual users, but adaptation may reinforce biases and narrow users’ information environments. Greater freedom to retrieve and formulate content expands relevance beyond user- and developer-set constraints.

  • Potential benefit: Personalized agent outcomes can be uniquely relevant for each user.
  • Risk: Adapting to user preferences can reinforce prejudices, create confirmation bias through selective retrieval, and establish echo chambers.
  • Application to agentic levels: Greater freedom to retrieve and formulate content increases the potential to provide information beyond user- and developer-set constraints.

D.0.3. VALUE: SUSTAINABILITY

AI agents may help address environmental problems, but their underlying models create environmental costs. Greater autonomy may also support novel solutions by combining more information and generating approaches beyond those anticipated by humans.

  • Potential benefit: AI agents may help address climate and efficiency problems through wildfire or flood forecasting and improved traffic efficiency.Such applications could decrease carbon emissions by helping address traffic inefficiency.
  • Risk: The models underlying current agents have environmental impacts including carbon emissions and potable-water use.
  • Application to agentic levels: Increasing autonomy may enable agents to harness more information and produce novel environmental solutions beyond those foreseen by humans.This potential coexists with environmental risks inherited from the models on which agents are based.

D.0.4. VALUE: TRUTHFULNESS

Truthfulness is a value AI agents may undermine as autonomy expands: systems can generate and amplify false information, while greater environmental control can separate agent-defined truth from human environments.

  • AI agents can generate false information, including deepfakes and misinformation, and use tailored, multi-platform output to manipulate beliefs and widen its impact.
  • As an agent gains control over its environment and resources, it can increasingly define what is true or false within that environment.
  • Because agent-created environments may differ from human environments, increased autonomy raises the potential for truthfulness to diverge from human-based standards.

E. Extended Methodology

The paper combines contemporary statements about AI agents with historical and influential literature to define agent autonomy, identify recurring value propositions, and analyze how values change with autonomy.

  • The authors collected 2024 and early-2025 statements from academic publications, online searches, industry surveys, blogs, case studies, and news articles.
  • They surveyed influential pre-2024 definitions, highly cited and relevant papers, and related concepts such as autonomous assistants and software agents.
  • The review uncovered a spectrum of AI-agent autonomy and developed a definition balancing definitional precision with recall while limiting vague or anthropomorphizing language.
  • The authors identified recurring AI-agent value propositions and analyzed how values are affected by increased autonomy.

F. Value Mapping

The paper maps AI-agent values and examples by compiling selected sources, systems, platforms, frameworks, pre-built agents, and articles that promoted agents as a major emerging technology.

  • Table 5 selects sources that reference different AI-agent values.
  • Table 6 distinguishes platforms for building agents, frameworks used to build them, and pre-built agent systems.
  • Table 7 selects online articles from 2024 that described AI agents as the “next big thing.”
Loading 2502.02649v3…