Source-linked AI summary

From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery

Jiaqi Wei, Yuejin Yang, Xiang Zhang, Yuhan Chen, Xiang Zhuang, Zhangyang Gao, Dongzhan Zhou, Guangshuai Wang, Zhiqiang Gao, Juntai Cao, Zijie Qiu, Ming Hu, Chenglong Ma, Shixiang Tang, Junjun He, Chunfeng Song, Xuming He, Qiang Zhang, Chenyu You, Shuangjia Zheng, Ning Ding, Wanli Ouyang, Nanqing Dong, Yu Cheng, Siqi Sun, Lei Bai, Bowen Zhou

arXiv:2508.14111v2cs.LG

TL;DR

Scientific discovery lacks a unified framework for understanding increasingly autonomous systems across processes, autonomy, and mechanisms. This survey integrates those perspectives into a domain-oriented synthesis of Agentic Science, identifying capabilities, workflows, applications, and challenges while highlighting memory and transparency as major boundaries.

  • Problem

    Existing surveys examine scientific workflows, autonomy scales, or agent architectures separately, leaving a unified framework for autonomous scientific discovery absent.

  • Method

    The survey connects foundational capabilities, core processes, and domain realizations while reviewing autonomous discovery across life sciences, chemistry, materials science, and physics.

  • Results

    The survey establishes a structured, domain-oriented synthesis of Agentic Science spanning autonomous reasoning, experimentation, analysis, iterative discovery, and applications across four natural-science domains.

  • Takeaways & Limitations

    Agentic Science is framed as a co-evolving research paradigm in which AI supports autonomous discovery while human inquiry remains part of the envisioned scientific system.

  • Takeaways & Limitations

    Scientific agents remain constrained by long-term causal reasoning and by memory systems’ difficulty maintaining high-fidelity, causally linked histories across extended research projects.

Abstract

from arXiv · show

Artificial intelligence (AI) is reshaping scientific discovery, evolving from specialized computational tools into autonomous research partners. We position Agentic Science as a pivotal stage within the broader AI for Science paradigm, where AI systems progress from partial assistance to full scientific agency. Enabled by large language models (LLMs), multimodal systems, and integrated research platforms, agentic AI shows capabilities in hypothesis generation, experimental design, execution, analysis, and iterative refinement -- behaviors once regarded as uniquely human. This survey provides a domain-oriented review of autonomous scientific discovery across life sciences, chemistry, materials science, and physics. We unify three previously fragmented perspectives -- process-oriented, autonomy-oriented, and mechanism-oriented -- through a comprehensive framework that connects foundational capabilities, core processes, and domain-specific realizations. Building on this framework, we (i) trace the evolution of AI for Science, (ii) identify five core capabilities underpinning scientific agency, (iii) model discovery as a dynamic four-stage workflow, (iv) review applications across the above domains, and (v) synthesize key challenges and future opportunities. This work establishes a domain-oriented synthesis of autonomous scientific discovery and positions Agentic Science as a structured paradigm for advancing AI-driven research.

1. Introduction

Agentic Science emerges as AI for Science advances from specialized computational tools toward autonomous, collaborative scientific inquiry. The survey unifies fragmented views of processes, autonomy, and mechanisms into a domain-oriented framework for autonomous discovery.

  • Agentic Science describes AI systems capable of formulating hypotheses, designing and executing experiments, interpreting results, and iteratively refining discovery.
  • Large language models enable scientific agents through natural-language understanding, complex reasoning, and tool use.
  • Existing surveys separately examine research processes, autonomy levels, or architectural mechanisms, leaving these perspectives fragmented.
  • The survey integrates foundational capabilities, core processes, and domain realizations across life sciences, chemistry, materials science, and physics.
  • It identifies five capabilities, models a dynamic four-stage workflow, reviews four natural-science domains, and synthesizes technical, ethical, and philosophical challenges.

2. The Evolution of AI for Science: From Tools to Autonomous Partners

AI for Science is described as evolving through distinct levels from specialized computational tools toward autonomous scientific partners. Agentic Science represents the emerging endpoint of this progression within the survey’s framework.

  • The evolution of AI in science progresses from narrowly scoped computational augmentation toward autonomous, end-to-end inquiry.

2.1. The Evolution of AI for Science

The evolution from tools to autonomous partners is organized into levels that progressively expand AI’s responsibility across scientific discovery. The framework culminates in open-ended agents and projects a future level involving invention of new scientific frameworks.

  • Level 1: AI as a Computational Oracle: Computational Oracles solve discrete, well-defined tasks but require human guidance for task definition, execution, and interpretation.
  • Level 2: AI as an Automated Research Assistant: Automated Research Assistants execute predefined workflow stages and tool sequences for well-defined subgoals, after which control returns to humans.
  • Level 3: AI as an Autonomous Scientific Partner: Autonomous Scientific Partners conduct the discovery cycle by forming hypotheses, running experiments, analyzing results, and iteratively refining knowledge with minimal human intervention.
  • Level 3: AI as an Autonomous Scientific Partner: Level 3 agents optimize cumulative scientific utility over an open-ended horizon while updating knowledge, evidence, hypotheses, and actions from new information.
  • Examples across domains: Examples span autonomous chemical reactions, therapeutic hypotheses, target discovery, nanobody design, chemical research, and materials discovery.
  • Level 4: AI as a Generative Architect: Generative Architects would invent scientific instruments, methodologies, conceptual models, or mathematical frameworks rather than work only within existing paradigms.

2.2. The Human Scientist’s Evolving Role

Agentic AI shifts scientists from executing research tasks toward setting goals, overseeing reliability and ethics, and validating outcomes. This human-agent arrangement also requires new skills and redistributes innovation across coordinated systems.

  • Scientists increasingly act as strategists who define research goals, maintain ethical and reliable methods, and integrate results into coherent narratives.
  • Human supervision includes constraining goals, checking reasoning, validating results against domain knowledge, and intervening when outputs drift.
  • Researchers must learn agent prompting, toolset management, and calibrated judgment about when to trust or scrutinize agent outputs.
  • Agents pursue subgoals within larger programs, exploring broad problem spaces and connecting ideas across disciplines in human-agent systems.

2.3. Agentic Science: The Focus of This Survey

Agentic Science is the survey’s focus, covering Level 2 task-level and Level 3 goal-level autonomy. The survey organizes this paradigm through capabilities, processes, and domain applications.

  • The survey’s framework links five foundational capabilities to four core processes in an agentic discovery loop.
  • These processes support analysis of autonomous scientific discovery across life sciences, chemistry, materials science, and physics.
  • Agentic Science encompasses Level 2 task-level autonomy and Level 3 goal-level autonomy in scientific research.Level 2 automates specific research tasks, whereas Level 3 independently pursues high-level scientific objectives.
  • Its defining principle is agency: purposeful, independent action within an environment to achieve a goal.

2.4. Modern Scientific Large Language Models for Agentic Science

Modern scientific large language models provide specialized and interactive capabilities for Agentic Science. The survey reviews their development strategies and applications across multiple scientific domains.

  • Scientific LLMs form a technological foundation for increasing levels of autonomy in scientific systems.The surveyed models extend beyond specialized computational-oracle roles toward more general reasoning and agentic capabilities.
  • Common development strategies include fine-tuning general-purpose models on curated scientific instruction data and domain-adaptive pre-training.
  • Domain-specific Sci-LLMs can improve specialized-task performance through curated datasets and subject-tailored training schemes.
  • Life Sciences: Life-science models support sequence representation learning, conversational analysis, cross-modal translation, and controllable biological-sequence generation.
  • Chemistry: Chemistry models integrate molecular structures, reaction data, and 3D conformations for property prediction, retrosynthesis, and drug design.
  • Materials Science: Materials-science models predict material properties and generate molecules or crystal structures with desired features.
  • Physics and Astronomy: Physics models combine language processing with physics engines and visual modules for parameter estimation, PDE solution operators, and knowledge retrieval.
  • Physics and Astronomy: Astronomy models adapt general architectures through continual pre-training and task-specific fine-tuning, with multimodal variants analyzing astronomical images.

3. Scientific Agents: Core Abilities and Challenges

Scientific agents require interconnected capabilities to manage long-horizon, empirically grounded discovery workflows. These capabilities support planning, tool use, memory, collaboration, and optimization, while introducing challenges involving reliability, multimodal data, coordination, and scientific validity.

  • Core abilities: Scientific agents combine five foundational pillars: planning and reasoning, tool integration, memory mechanisms, multi-agent collaboration, and optimization and evolution.Together, these capabilities support long-term, iterative, empirically grounded research workflows rather than discrete short-horizon tasks.
  • Planning and reasoning: Planning and reasoning engines translate scientific goals into executable actions, including hypothesis formulation, experiment design, code execution, and database queries.Linear chains such as plan-and-solve and Chain-of-Thought provide basic decomposition, while tree-based methods explore alternatives and backtrack from unpromising paths.
  • Tool use and integration: Tool integration extends agents beyond language-model limits in computation, data access, and physical-world interaction, spanning foundational utilities, domain tools, and experimental or simulation platforms.Scientific workflows require interoperable chains of specialized tools, with precise parameterization and detailed provenance records.
  • Memory mechanisms: Memory supports information retention, learning from experience, contextual continuity, iterative refinement, knowledge accumulation, and hypothesis testing.The paper distinguishes memory for iterative task execution from memory as a knowledge hub connected to external information repositories.
  • Collaboration between agents: Multi-agent collaboration distributes tasks, synthesizes diverse information, and iteratively refines solutions through structured workflows, deliberative refinement, and dynamic adaptation.Scientific collaboration must coordinate heterogeneous data, specialized software, and physical instruments while preserving empirical validity and epistemic diversity.
  • Optimization and evolution: Scientific-agent optimization is constrained by expensive evaluations, sparse rewards, physical reality, safety protocols, and the need for reproducible, verifiable knowledge.Scientific feedback may require days of laboratory work, while rare breakthroughs make meaningful policy learning difficult.

4. Agentic Science: Dynamic Workflow and Challenges

Agentic Science frames discovery as a dynamic, autonomous closed loop spanning hypothesis generation, experimentation, analysis, and iterative evolution. Its implementation combines memory, reasoning, tool use, collaboration, and self-improvement, while facing challenges from heterogeneous data, noisy feedback, and long-horizon causal reasoning.

  • Workflow overview: The workflow comprises four stages: observation and hypothesis generation; experimental planning and execution; result analysis; and synthesis, validation, and evolution.Execution order may be dynamically adjusted according to agent objectives, context, and ongoing results.
  • Observation and Hypothesis Generation: Hypothesis generation connects retrieved and structured scientific knowledge to exploratory reasoning over candidate research directions.Agents use techniques such as RAG and knowledge graphs before formulating hypotheses through pattern discovery and symbolic reasoning.
  • Observation and Hypothesis Generation: Agentic systems have generated biologically meaningful hypotheses, including validated cancer targets, a repurposed dry-eye drug candidate, and a link between transcriptional noise and brain aging.OriGene’s GPR160 and ARG2 targets were subsequently validated in patient-derived systems, while Robin and CellVoyager surfaced additional findings.
  • Experimental Planning and Execution: Experimental execution maps abstract plans to tool invocations, code, or robotic hardware, enabling end-to-end laboratory and computational research loops.Examples include ORGANA’s 19-step synthesis protocol, which reduced human workload by over 80%, and a virtual pipeline that designed 92 nanobodies, two of which showed strong binding.
  • Result Analysis: Result analysis combines multimodal data extraction with structured reasoning, but noisy feedback and heterogeneous formats require robustness against signal confusion and confirmation bias.Agents must integrate text, tables, genomic sequences, and imagery when interpreting experimental outcomes.
  • Synthesis, Validation, and Evolution: Iterative evolution uses reflection, recursive evaluation, feedback, and experiment-guided ranking to improve hypotheses, protocols, and analysis pipelines across research loops.Long-term progress remains constrained by the difficulty of maintaining high-fidelity, causally linked histories over months or years; multi-agent collaboration also supports broader research pipelines.

5. Agentic Life Sciences Research

Agentic AI is being applied across life sciences to automate complex analyses, experimental design, therapeutic discovery, and protein engineering. The reviewed systems range from specialized workflow assistants to agents that execute iterative discovery cycles with limited human intervention.

  • Scope: Life-science agents address complex, multi-step workflows spanning omics analysis, hypothesis generation, experimental design, and result interpretation.Applications cover genomics, transcriptomics, proteomics, drug discovery, and protein engineering.
  • General frameworks: General biomedical frameworks emphasize modular tools, self-evolution, and multi-agent collaboration to support diverse research tasks.STELLA expands its toolset through an evolving Template Library and dynamic Tool Ocean, while Biomni decomposes queries into tool-based plans.
  • Omics analysis: Single-cell and multi-omics agents automate pipelines from data processing through workflow design, code generation, evaluation, and reporting.CellAgent uses planner, executor, and evaluator agents with iterative optimization of tools and hyperparameters.
  • Protein science: Protein-design systems combine specialized agents, structure prediction, simulations, and iterative reflection to generate proteins with targeted properties.ProtAgents coordinates retrieval, structure analysis, and physics-based simulation, while Sparks executes hypothesis generation, experiment design, and refinement.
  • Drug discovery: Drug-discovery agents support end-to-end pipelines and targeted discovery stages, producing pharmaceutical candidates, therapeutic targets, and experimentally validated hypotheses.LIDDiA generated molecules meeting key criteria for over 70% of 30 clinically relevant targets; OriGene nominated GPR160 and ARG2, and Robin identified ripasudil for dAMD.
  • Specialized applications: Specialized agents also improve experimental design and therapeutic reasoning, including genetic perturbation prediction and nanobody generation.BioDiscoveryAgent improved relevant genetic-perturbation prediction by 21% on average, while a human-guided pipeline produced 92 nanobody candidates, two of which showed improved binding.

6. Agentic Chemistry Research

Agentic chemistry research combines language models with chemical tools, specialized agents, and robotics to automate reasoning, synthesis, simulation, and experimental execution. The reviewed systems span general-purpose frameworks, reaction optimization, robotic assistance, molecular design, and quantum-chemistry workflows.

  • Scope: Chemical agents target the full research workflow, from hypothesis generation and planning to experiment execution, analysis, and validation.The survey emphasizes tool integration, robust reasoning, literature comprehension, and robotic platforms as foundations for autonomy.
  • General frameworks: General-purpose systems coordinate specialized agents and external tools for multi-step chemical research with reduced human intervention.ChemCrow uses 18 expert-designed tools, while ChemAgents coordinates literature reading, experiment design, computation, and robot operation.
  • Organic synthesis: Reaction-optimization systems automate experimental planning, robotic execution, monitoring, and interpretation across complete synthesis workflows.Coscientist optimized palladium-catalyzed cross-couplings, while LLM-RDF coordinated literature review, screening, scale-up, and purification.
  • Robotic experimentation: Robotic assistants translate natural-language goals into physical laboratory procedures and can reduce researchers’ physical workload and time demands.ORGANA automates solubility testing, pH measurement, and recrystallization; user studies reported over 50% lower physical demand and an average 80% time saving.
  • Molecular design: Generative chemistry systems combine language models, diffusion or predictive models, quantum validation, and feasibility constraints to discover novel structures.MOFGen generated hundreds of thousands of candidate MOFs and experimentally synthesized five new structures.
  • Simulation: Agentic quantum-chemistry assistants generate and execute simulation workflows from natural-language prompts using hierarchical memory and adaptive tool selection.El Agente Q achieved an average success rate of over 87% on benchmark tasks.

7. Agentic Materials Science Research

Agentic materials-science systems address vast design spaces and complex workflows through knowledge-grounded retrieval, human-in-the-loop planning, hypothesis generation, material discovery, and simulation automation. Their applications include novel materials, inverse design, specialized material classes, and laboratory characterization.

  • Scope: Materials-science agents are organized around novel-material discovery, simulation and characterization automation, and general discovery platforms.The field’s main challenges include integrating heterogeneous data, reliable knowledge grounding, autonomous planning, and human-AI collaboration.
  • General methodologies: Knowledge-grounded platforms use retrieval, databases, and multimodal agents to improve reliability and unify heterogeneous materials information.LLaMP retrieves materials data and runs simulations, while multicrossmodal agents combine image, text, table, and video information in a shared embedding space.
  • Discovery platforms: Human-in-the-loop platforms generate hypotheses, plan experiments, invoke physics-based tools, and iteratively learn from feedback.MatPilot closes the loop between hypotheses, experiments, and optimization; MAPPS includes workflow planning, code generation, and scientific mediation.
  • Design and discovery: Materials-discovery agents navigate combinatorial design spaces by generating hypotheses and refining candidates using physical or computational feedback.SciAgents uses a knowledge graph to generate and refine hypotheses, while TopoMAS combines retrieval, first-principles validation, and continuously updated knowledge.
  • Advanced materials: Specialized systems discover topological materials and automate metamaterial design through coordinated generative, simulation, and reasoning agents.TopoMAS identified the novel topological phase SrSbO3; CrossMatAgent produces simulation-ready metamaterial designs.
  • Inverse design: Inverse-design agents propose structures for target properties and support autonomous or human-guided evaluation with surrogate or computational models.dZiner proposes compounds from literature insights and iteratively evaluates them with surrogate models.
  • Simulation and characterization: Natural-language interfaces make complex simulations and characterization workflows more accessible while iterative correction and benchmarking remain important.ChemGraph supports DFT and other atomistic calculations, Foam-Agent automates CFD workflows, and AILA shows multi-agent architectures outperform single agents in AFM automation.

8. Agentic Physics and Astronomy Research

Agentic AI is being applied across physics, astronomy, quantum computing, and engineering to automate research workflows, while current systems still face reasoning and mission-complexity limitations.

  • Agentic systems assist physics and astronomy research by managing software, analyzing data, formulating and testing hypotheses, and automating experimental procedures.
  • General Frameworks and Methodologies: MoRA improved open-source LLM accuracy on SciEval and PhysicsQA by up to 16%, while LP-COMDA reduced power-converter design error by 63.2% and ran over 33 times faster than human-led processes.
  • Astronomy: StarWhisper automates the NGSS observational workflow, including customized observation lists, telescope operations, real-time transient detection, and follow-up proposal generation.
  • Astronomical Research: AI Cosmologist and related multi-agent systems automate stages from idea generation and experimental design through analysis and scientific publication production.
  • Computational Mechanics and Fluid Dynamics: OpenFOAMGPT 2.0 achieved 100% success and reproducibility across over 450 test cases through specialized agents for preprocessing, prompting, simulation, and post-processing.
  • Quantum Computing: The k-agents framework converts laboratory knowledge and experimental procedures into agent-based state machines whose analyzed results drive closed-loop quantum-experiment transitions.

9. Challenges in Agentic Science

Agentic Science introduces challenges for reproducibility, novelty validation, interpretability, accountability, and risk governance that extend beyond conventional LLM limitations.

  • Agentic Reproducibility and Reliability: Agentic discovery trajectories are stochastic and context-sensitive, making scientific findings difficult to reproduce consistently.
  • Agentic Reproducibility and Reliability: Execution accuracy can be as low as 39%, while catastrophic forgetting and prompt sensitivity undermine stable knowledge and reproducible multi-stage outcomes.
  • Requirements: Reproducibility, interpretability, and ethical governance must be treated as integral design considerations for trustworthy agentic scientific systems.
  • Validation of Novelty: Novel hypotheses are difficult to distinguish from interpolation or hallucination because LLMs may produce repetitive, plausible-sounding, false, or unverifiable content.
  • Interpretability and Trust: Model opacity prevents reliable auditing of reasoning lineage and undermines validation, trust, and assimilation of AI-generated scientific insights.
  • Ethical and Societal Risks: Autonomous agents raise accountability and dual-use risks, including responsibility for erroneous findings or hazardous discoveries and misuse involving toxins, pathogens, or harmful technologies.

10. Future Outlook of Agentic Science

The future outlook extends Agentic Science beyond workflow automation toward autonomous invention, interdisciplinary synthesis, global agent cooperation, and discovery-level benchmarks.

  • From Automation to Autonomous Invention: Autonomous invention would shift agents from tool users to tool creators capable of proposing novel instruments or conceptual frameworks.
  • Interdisciplinary Synthesis: Future multimodal agents could support interdisciplinary synthesis by surfacing analogies between previously siloed scientific domains.
  • Global Cooperation: A decentralized ecosystem could connect specialized agents across institutions for peer review, hypothesis refinement, and collaborative experimentation at planetary scale.
  • Requirements: Realizing these futures requires advances in secure federated protocols, multi-agent reasoning, interpretability, accountability, and scientific validation.
  • The Nobel-Turing Test: The Nobel-Turing Test asks whether an autonomous agent or hybrid team can identify foundational gaps, generate testable hypotheses, design novel methods, and produce Nobel-worthy discoveries.

11. Conclusion

Agentic Science is presented as a structured stage in AI for Science that connects autonomous reasoning, experimentation, and iterative discovery across multiple natural-science domains.

  • The survey synthesizes agentic discovery across life sciences, chemistry, materials science, and physics through a framework linking capabilities, processes, and domain realizations.
  • It presents Agentic Science as augmenting human inquiry while highlighting technical, ethical, and philosophical challenges for trustworthy progress.
Loading 2508.14111v2…