Source-linked AI summary

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems

Leon Staufer, Kevin Feng, Kevin Wei, Luke Bailey, Yawen Duan, Mick Yang, A. Pinar Ozisik, Stephen Casper, Noam Kolt

arXiv:2602.17753v2cs.CYcs.AI

TL;DR

Agentic AI is advancing while its development and deployment remain difficult to track because public documentation is inconsistent. The paper constructs the 2025 AI Agent Index from 30 prominent deployed agents and finds substantial transparency gaps, especially around safety and evaluation. Its scope is limited by uneven reporting and selection criteria that favor significant, publicly available systems.

  • Problem

    Public information is insufficient to answer basic questions about agent developers, deployment, evaluation, and safeguards, creating a documentation gap for researchers and policymakers.

  • Method

    The paper systematically selects and annotates 30 prominent deployed agentic systems using revised criteria and information fields.

  • Results

    Most safety-related fields lack public information, and only four indexed agents provide agent-specific safety evaluations.

  • Takeaways & Limitations

    The Index documents ecosystem-wide differences in transparency and identifies persistent gaps in safety, evaluation, and ecosystem reporting.

  • Takeaways & Limitations

    The Index favors significant, publicly available, general-purpose agents, while smaller, emerging, domain-specific, and closed-door systems may show different patterns.

Abstract

from arXiv · show

Agentic AI systems are increasingly capable of performing professional and personal tasks with limited human involvement. However, tracking these developments is difficult because the AI agent ecosystem is complex, rapidly evolving, and inconsistently documented, posing obstacles to both researchers and policymakers. To address these challenges, this paper presents the 2025 AI Agent Index. The Index documents information regarding the origins, design, capabilities, ecosystem, and safety features of 30 state-of-the-art AI agents based on publicly available information and email correspondence with developers. In addition to documenting information about individual agents, the Index illuminates broader trends in the development of agents, their capabilities, and the level of transparency of developers. Notably, we find different transparency levels among agent developers and observe that most developers share little information about safety, evaluations, and societal impacts. The 2025 AI Agent Index is available online at https://aiagentindex.mit.edu

1 Introduction

The 2025 AI Agent Index addresses limited public information about impactful agentic systems by documenting 30 widely deployed agents and analyzing ecosystem-wide transparency patterns.

  • Public information remains limited about who develops impactful agents, where they are deployed, how they are evaluated, and which guardrails they use.
  • The Index documents 30 agentic systems across legal, technical capabilities, autonomy and control, ecosystem interaction, evaluation, and safety categories.
  • The 2025 Index revises its inclusion criteria and information fields to reflect recent growth and change in the agent ecosystem.
  • Most safety-related fields, 135/240, lack public information; nearly all indexed agents rely on GPT, Claude, or Gemini foundation-model families.
  • Only four agents provide agent-specific safety evaluations, and most agents do not disclose their AI nature to end users or third parties by default.
  • The paper contributes the Index, ecosystem-wide trend analysis, and case studies of browser agents, agentic chatbots, and customizable enterprise-agent builders.

2 Background and Related Work

Prior work describes agents, their growing deployment and risks, and documentation frameworks, but lacks a comparable framework specifically for documenting agentic AI systems.

  • Definitions of AI agents vary across disciplines but commonly emphasize autonomy, goal-directedness, and complex long-horizon task completion.
  • Research and enterprise interest in AI agents has increased sharply, with 2025 papers mentioning agents exceeding the total from 2020–2024 combined by more than twofold.
  • Agents can create new risks because they may act directly in the world, including through autonomous actions such as hacking websites.
  • Related work includes agent evaluations, landscape lists, capability benchmarks, operational-visibility studies, and economic and governance analyses.
  • Existing documentation frameworks cover models, systems, evaluations, usage, safety, and ecosystems, but comparable agent-specific frameworks were absent aside from prior AI Agent Indexes.
  • The 2025 Index revises the 2024 Index by examining fewer systems in greater depth and separating chat, browser, and enterprise-agent categories.

3 Constructing the 2025 AI Agent Index

The Index was constructed by systematically selecting deployed, impactful, general-purpose agents using agency, impact, and practicality criteria, then annotating them across six information categories.

  • Inclusion criteria: Included systems had to satisfy all agency criteria, at least one impact criterion, and all practicality criteria as of December 31, 2025.
  • Agency criteria: Agency required autonomy, complex goals, environmental interaction through tools or APIs, and generality across under-specified tasks.
  • Agency criteria: Autonomy was operationalized as at least intermediate autonomy, where agents perform most tasks independently but rely on users for critical determinations.
  • Impact criteria: The Index used impact criteria including public interest, market significance, and developer significance, with at least one required for inclusion.
  • Practicality criteria: Practicality required public availability, off-the-shelf deployability with minimal configuration, and general-purpose task capability.
  • Agent categories: The 30 systems were grouped into chat applications with agentic tools, browser-based agents, and enterprise workflow agents with distinct interfaces and governance challenges.
  • Selection process: Researchers identified 95 candidate agents, screened them against the criteria, consulted Chinese ecosystem experts, and cross-referenced multiple agent databases.
  • Annotation: Agents were annotated across six categories, producing 45 information fields per system covering capabilities, control, ecosystem interaction, safety, evaluation, and accountability.

4 Findings

The Index maps 30 agents across six dimensions and reveals recent deployment growth, diverse interfaces and autonomy levels, concentrated development, and substantial variation in interoperability, control, and transparency.

  • The Index covers 30 agents across six categories spanning product overview, accountability, technical architecture, autonomy and control, ecosystem interaction, and safety, evaluation, and impact.
  • 24/30 agents were released or received major agentic feature updates during 2024–2025, indicating a recent surge in deployment.
  • Company and accountability: Developers are concentrated in the United States and China: 13/30 agents come from Delaware-incorporated companies, 5/30 from China-incorporated companies, and 4/30 from other incorporations.
  • Safety and accountability: 15/30 agents reference AI safety frameworks, while 10/30 have no safety framework documentation; enterprise assurance standards are more widely adopted.

5 Illustrative Case Studies

The case studies show how chat, browser, and enterprise agents differ in autonomy, safety practices, ecosystem dependencies, and transparency. They also expose distinct governance challenges arising from approval gates, untrusted web content, and shared platform control.

  • Case-study scope: The three case studies cover ChatGPT Agent, Perplexity Comet, and HubSpot Breeze Agents as distinct chat, browser, and enterprise agent categories.They were selected as reference points for understanding common and divergent features across categories.
  • ChatGPT Agent: ChatGPT Agent combines L2-L4 autonomy with approval for sensitive operations and a hosted virtual computer with restricted network access.Its evaluation covers policy compliance, jailbreaks, hallucinations, fairness, CBRN risks, cyber capabilities, and autonomy.
  • ChatGPT Agent: ChatGPT Agent is the only indexed system using cryptographic HTTP-request signing, addressing identity and auditability challenges affecting browser agents.It is also one of two systems with a dedicated agent-specific system card.
  • Perplexity Comet: Perplexity Comet operates at L4-L5 autonomy and proceeds autonomously after initiation, but no agent-specific evaluations, third-party testing, or benchmark results were disclosed.Security researchers identified indirect prompt injection and URL-based attacks against browser agents processing untrusted web content.
  • Perplexity Comet: Browser agents operate across third-party environments without established mechanisms to negotiate or verify interaction terms, while Comet’s autonomy heightens these trust challenges.Existing protocols such as robots.txt were designed for crawlers rather than autonomous actors.
  • HubSpot Breeze Agents: Breeze uses split autonomy, ranging from L1-L3 during design to L5 after deployment, with approval configurable during creation and automatic triggers able to run without approval.Its action space centers on internal databases and organizational tools, while constraints rely on tool permissions and user roles.
  • HubSpot Breeze Agents: Enterprise agents are joint products of platforms and business users, leaving neither side with complete control or visibility over the other’s contribution.Breeze emphasizes compliance, while agent-specific guardrails become the user’s responsibility; its safety disclosures provide no testing methodology or results.

6 Discussion

The discussion identifies persistent transparency gaps, concentrated dependence on a few foundation-model families, and fragmented accountability across the agent ecosystem. It argues that deployment-specific information and stronger governance mechanisms are needed, while noting important scope and reporting limitations.

  • Key findings: The Index covers 1,350 fields for 30 agents and finds persistent reporting limitations, especially for ecosystem and safety-related features.The authors state that existing transparency expectations are largely unmet.
  • Transparency: Only ChatGPT Agent, OpenAI Codex, Claude Code, and Gemini 2.5 Computer Use provide agent-specific system cards, while 9/30 agents report capability-benchmark performance.The Index characterizes this as inconsistent and selective reporting, particularly around safety.
  • Ecosystem concentration: Almost all indexed systems rely on GPT-, Claude-, or Gemini-family models, creating potential single points of failure through pricing changes, outages, and safety regressions.Only US- and China-based foundation-model developers in the Index operate proprietary models.
  • Accountability: Reliable agent evaluation is difficult because models, orchestration platforms, builders, and deployments form dependency chains with context-specific tools and autonomy levels.The authors state that no single entity bears clear responsibility across these layers.
  • Web governance: Browser agents often bypass robots.txt, shifting control away from content hosts and creating tension with established web-scraping norms.The discussion identifies allowlisting and cryptographic authentication as potential alternative governance mechanisms.
  • Web governance: Without cryptographic signing, verifying agent identity or proving what an agent did becomes significantly harder, while disclosure responsibilities may shift from developers to operators.This raises questions about whether end users will know when they are interacting with AI.
  • Limitations and outlook: The Index favors significant, publicly visible agents and English- and Mandarin-language documentation, which may limit generalizability and understate practices documented in other languages.It also excludes domain-specific agents and may miss internal evaluations or risk-management practices unavailable publicly.
  • Limitations and outlook: Future work should extend coverage to internal and domain-specific agents and track transparency and governance patterns as the ecosystem evolves.The Index is intended as a baseline for measuring future transparency improvements or regressions.

A.1 Further Analysis

The further analysis tracks the growth, release timing, information completeness, and category-specific support features of indexed agents. It highlights both rapid productization and uneven documentation across categories.

  • Release trends: Figure 9 tracks the first release of indexed agentic products over time and across chat, enterprise, and browser categories.The figure provides the temporal and categorical layout for comparing product emergence.
  • Information availability: 227 of 1,350 information fields contain no information, most often in Ecosystem Interaction and Safety, Evaluation, and Impact categories.Non-empty information fields average 14 words.
  • Release trends: Figure 11 reports the number of new AI agentic product releases by month.Monthly release counts show the timing of product deployment activity.
  • Information availability: Figure 12 counts fields marked “None found” by field category, emphasizing category-level differences in publicly available information.The figure focuses specifically on missing-information counts rather than product releases.
  • Category differences: Enterprise agents more often support model selection and MCP, with support in 9/13 and 12/13 systems respectively.Figure 13 compares these capabilities by agent category.

A.2 Sample Entry of the Index: Claude Code

The Claude Code sample entry documents its identity, deployment context, technical architecture, autonomy controls, ecosystem behavior, and safety documentation. It presents a permission-based coding agent with sandboxing, monitoring, evaluations, and reported vulnerabilities.

  • Technical Capabilities & System Architecture: The agent operates over filesystems, bash commands, and MCP through a terminal chatbot interface, with hierarchical markdown memory and selectable Claude models.Its component accessibility is documented as closed source.
  • Autonomy & Control: Claude Code ranges from chatbot-like planning to multistep tool use, while requiring permission for bash commands, file edits, or files outside its initial directory.Users can pause or stop the agent, and summarized execution traces and context-use statistics are visible.
  • Ecosystem Interaction: Anthropic identifies Claude-related web activity through ClaudeBot, Claude-User, and Claude-SearchBot user-agent tokens and supports MCP as an interoperability standard.Anthropic states that these bots respect robots.txt directives and anti-circumvention technologies, while independent accounts report contrary behavior during some periods.
  • Safety, Evaluation & Impact: Claude Code uses permission controls, content classifiers, site blocklists, action confirmations, and sandboxing to constrain potentially harmful actions.Its documentation also references agentic-misuse evaluations, third-party red-teaming, bug-bounty support, and an AI-orchestrated cyber-espionage incident.

B Annotation Methodology

The annotation methodology began with a broad product list and focused the final Index on agentic products supporting general customization. The considered products span coding, enterprise, browser, chatbot, and other agent categories.

  • Product Selection: The study distinguishes general-purpose agentic products from ready-made agents embedded in broader enterprise products.The analysis focuses on products allowing general customization and on agents that can be created through them.
  • Product Selection: The considered product list includes commercial and open-source systems from developers such as Anthropic, Google, Microsoft, OpenAI, Salesforce, and many others.The list covers coding agents, enterprise platforms, browser agents, chatbot products, and agent-building systems.
  • Annotation Fields: The methodology records product-level metadata including public interest, developer identity, legal structure, governance documents, and safety or accountability information.The fields include search volume, GitHub stars, market capitalization or valuation, and developer governance characteristics.

4. Technical Capabilities & System Architecture

The Index characterizes agent technical capabilities across models, interfaces, memory, tools, autonomy, user approval, identification, interoperability, and web conduct. These fields describe both what agents can do and how users or external systems interact with them.

  • Technical Capabilities: Model and architecture fields cover selectable models, documentation, observation space, action space, sandboxing, and tools with write access.Observation sources include internet access and MCP support, while action fields assess whether agents can affect the real world without human approval.
  • Interaction Design: The Index records memory architecture, interface design, anthropomorphism warnings, human–agent teaming practices, and user roles.User roles include designing agents, operating them, acting on outputs, and examining their behavior.
  • Autonomy & Control: Autonomy is classified from L1 to L5, ranging from on-demand user support to systems with no user involvement.The Index also records planning depth, approval requirements, monitoring, and emergency-stop mechanisms.
  • Ecosystem Interaction: Ecosystem fields examine whether agents identify themselves to humans, provenance tracking, digital signatures, web-request identifiers, interoperability standards, and integrations.Examples of interoperability standards include AGNTCY, Agent Connect Protocol, and Model Context Protocol.

7. Safety, Evaluation & Impact

The safety and evaluation methodology catalogs guardrails, containment, evaluations, incidents, and documentation practices for agents. It combines structured annotation, public-source research, automated verification, manual checking, and developer correction requests.

  • Safety Fields: Safety fields cover technical guardrails, sandboxing and containment, evaluated risks, benchmark performance, third-party testing, vulnerability disclosure, and known incidents.The methodology distinguishes built-in protections from optional guardrails available on agent-building platforms.
  • Annotation Principles: Annotators were instructed to document object-level findings, agent-specific safety and transparency features, platform defaults, sources, and uncertainty markers.They were told to distinguish agent evaluations from evaluations of underlying foundation models.
  • Information Sources and Scope: The study relied on official documentation, company materials, trust centers, conference demonstrations, and publicly visible interfaces, without conducting experiments or analyzing open-source code.For open-source agents, annotators used documentation and README files rather than code analysis.
  • Candidate Identification: LLM-based research queries generated an initial candidate list, after which researchers screened candidates against inclusion criteria and consulted Chinese ecosystem experts.The process also cross-referenced prior indexes and agent lists.
  • Verification: The verification pipeline used web research, structured JSON verification, and manual comparison of original and LLM-generated annotations.Annotators manually checked sources, amended results when needed, and used a custom viewer for side-by-side comparison.

C Public Interest Methodology

The methodology generates product-specific public-interest search terms in two stages: researching each product, then producing new, concise, natural, and unambiguous queries.

  • Search-volume metrics combine product and company names with simpler product-only or company-plus-agent terms when unambiguous.
  • GPT-5.2 with web search first produces concise product summaries covering identity, use case, key features, and distinguishing characteristics.
  • GPT-5.2 then generates new search terms from the product background, company, product name, context, and existing terms.
  • Terms must be specific, unambiguous, concise, natural, and directly tied to the product rather than the company generally.
  • Only new terms are retained, with duplicates excluded and an empty array returned when no suitable new terms exist.
  • The procedure includes company-product combinations, reversed order when applicable, standalone product names when unambiguous, and company-plus-category queries.

C.2 Public Interest Analysis

The public-interest analysis presents search-volume trends, GitHub engagement, and documented safety information for indexed AI agents through five figures.

  • Figure 16 shows global search volume for each AI agent product on a logarithmic scale and marks products included in the Index.
  • Figure 17 shows monthly logarithmic search volume for each indexed product’s most popular term during 2025, grouped by chat, enterprise, and browser agents.
  • Figure 18 compares monthly logarithmic search volume for the 10 most popular AI agent products by agent category.
  • Figure 19 reports GitHub stars for repositories associated with the agent products.
  • Figure 20 illustrates limited details on agent safety features and red-teaming methodology using HubSpot’s Breeze Agents model card.
Loading 2602.17753v2…