Source-linked AI summary

Artificial Intelligence Index Report 2026

Sha Sajadieh, Loredana Fattorini, Raymond Perrault, Yolanda Gil, Vanessa Parli, Lapo Santarlasci, Juan Pava, Nestor Maslej, Russ Altman, Erik Brynjolfsson, Carla Brodley, Jack Clark, Virginia Dignum, Vipin Kumar, James Landay, Terah Lyons, James Manyika, Juan Carlos Niebles, Yoav Shoham, Elham Tabassi, Russell Wald, Toby Walsh, Dan Weld

arXiv:2606.15708v3cs.AI

TL;DR

AI systems are advancing faster than the governance, evaluation, education, and data systems needed to keep pace. This report tracks that widening gap across testing, economic and labor effects, sovereignty, science, and medicine, finding that AI’s growing capabilities increasingly challenge existing systems for measuring and managing its impact.

  • Problem

    Governance, evaluation, education, and data systems are struggling to keep pace with rapidly advancing AI, limiting preparedness to measure and manage its impact.

  • Method

    The report tracks AI testing across reasoning, safety, and real-world task execution while adding analyses of economic value, labor effects, sovereignty, science, and medicine.

  • Results

    The report finds that measurements of AI capabilities are becoming increasingly difficult to rely on as systems are tested more ambitiously across multiple domains.

  • Takeaways & Limitations

    The report frames the widening gap between AI capabilities and the systems needed to evaluate and govern them as a central challenge across its chapters.

  • Takeaways & Limitations

    The government AI spending analysis may not fully represent all countries or regions because data availability and quality vary significantly.

Abstract

from arXiv · show

Welcome to the ninth edition of the AI Index report. As AI continues to advance rapidly, the question becomes whether the systems built around it can keep up. Governance frameworks, evaluation methods, education systems, and the data infrastructure needed to track AI's impact are struggling to match the pace of the technology itself. That gap between what AI can do and how prepared we are to manage it runs through every chapter of this year's report. New in this edition, the report tracks how AI is being tested more ambitiously across reasoning, safety, and real-world task execution, and why those measurements are increasingly difficult to rely on. It also features new estimates of generative AI's economic value alongside emerging evidence of its labor market effects, an analytical framework on AI sovereignty, and a science chapter developed in collaboration with Schmidt Sciences. For the first time, the report features standalone chapters on AI in science and AI in medicine, reflecting AI's growing impact across these two domains.

1 Research and Development … Appendix

The ninth AI Index report examines whether governance, evaluation, education, and data systems can keep pace with rapidly advancing AI. It documents expanding adoption and deployment while emphasizing measurement gaps, benchmark saturation, and uneven policy responses.

  • Appendix: The report tracks AI’s rapid advance against systems for governance, evaluation, education, and impact measurement that struggle to keep pace.This gap between capability and preparedness runs through every chapter of the report.
  • Appendix: The AI Index provides independently sourced global data for decision-makers while identifying where the available picture remains incomplete.Its stated purpose is to support informed decisions as AI reshapes work, learning, classrooms, clinics, and legislatures.
  • Appendix: Generative AI reached nearly 53% population-level adoption within three years, while organizational adoption rose to 88%.Global corporate investment more than doubled in 2025, and leading AI companies reached meaningful revenue scale faster than previous technology generations.
  • Appendix: Frontier models are converging and open-weight models are becoming more competitive, while benchmarks saturate and independent testing does not always confirm developer reports.Frontier labs are also disclosing less, making evaluation increasingly difficult to rely on.
  • Appendix: AI shifted in science from accelerating individual research steps toward attempting to replace entire workflows, while clinical AI moved from pilots toward broader deployment.Ambient AI scribes are scaling across health systems.
  • Appendix: Governments adopted divergent AI policies in 2025: the EU AI Act’s first prohibitions took effect, while the United States shifted toward deregulation.Japan, South Korea, and Italy passed national AI laws, and more than half of newly adopted national AI strategies came from developing countries.
  • Appendix: AI sovereignty emerged as a central organizing principle across governments’ policy efforts.The report presents these responses as varied rather than pointing in a single direction.

CHAIR CO-CHAIR

The report’s chair/co-chair affiliation is the University of Southern California’s Information Sciences Institute.

  • CHAIR CO-CHAIR: The chair/co-chair is affiliated with the University of Southern California, Information Sciences Institute.

MEMBERS

The members represent affiliations spanning JPMorgan Chase & Co., Google, the University of Oxford, Stanford University, Salesforce, and AI21 Labs.

  • JPMorgan Chase & Co. is listed as a member.
  • Google and the University of Oxford are listed together as members.
  • Stanford University and Salesforce are listed together as members.
  • Stanford University and AI21 Labs are also listed together as members.

LEAD AND EDITOR-IN-CHIEF RESEARCH MANAGER … Ipsos

AI capability, infrastructure, adoption, and scientific applications expanded rapidly in 2025, while transparency, responsible-AI evaluation, education, talent mobility, and governance struggled to keep pace. The report also finds that AI’s benefits and risks are unevenly distributed across countries, sectors, physical environments, and scientific and clinical tasks.

  • Technical Performance: AI performance reached or exceeded human baselines in advanced science, multimodal reasoning, mathematics, and coding, but remained jagged: agents reached ~66% OSWorld success, top models read analog clocks correctly only 50.1% of the time, and robots completed 12% of household tasks.Controlled robotic manipulation reached 89.4% on RLBench, illustrating the gap between predictable simulations and unpredictable household environments.
  • Responsible AI: Responsible-AI measurement lagged capability progress: documented incidents rose from 233 to 362, leading developers reported capability results far more consistently than responsible-AI results, and improving safety could degrade accuracy.The report frames benchmark reliability and AI-system evaluation as increasingly difficult as systems become more capable.
  • Economy and Labor: The United States led private AI investment with $285.9 billion—more than 23 times China’s $12.4 billion—but its inflow of AI researchers and developers fell 89% since 2017, while productivity gains coincided with nearly 20% lower employment among U.S. developers aged 22 to 25.Generative AI reached 53% population adoption within three years, and its estimated annual value to U.S. consumers reached $172 billion by early 2026.
  • Science and Medicine: AI showed substantial but uneven scientific and clinical value: frontier models outperformed human chemists on average yet scored below 20% on astrophysics replication, while clinical-note tools reduced physician writing time by up to 83% despite limited rigorous evidence.Smaller specialized models also surpassed much larger systems on selected protein and genomics benchmarks.
  • Research and Development: Industry produced 91.2% of notable AI models in 2025, while the United States released 59 models to China’s 35 and frontier-model disclosures increasingly withheld training code, data, and parameters.Training-code access narrowed sharply: 81 of 102 notable models lacked corresponding training code, while only 4 were open source.
  • AI Infrastructure: Global AI compute capacity grew 3.3x annually to 17.1 million H100-equivalents, but infrastructure remained geographically and industrially concentrated, with 5,427 U.S. data centers and TSMC fabricating almost every leading AI chip.Nvidia supplied over 60% of total compute, while the United States hosted more than ten times as many data centers as any other country.
  • Environmental Impact: AI’s environmental footprint expanded alongside deployment: data-center power capacity reached 29.6 GW, Grok 4 training emissions reached 72,816 tons of CO2 equivalent, and annual GPT-4o inference water use may exceed the drinking-water needs of 1.2 million people.Hardware became about ten times more efficient per watt over a decade, but model scaling caused total training power and emissions to continue rising.
Loading 2606.15708v3…