Source-linked AI summary
Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions
Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes, Christina Mack
TL;DR
Agentic AI promises to automate increasingly complex scientific workflows, but existing systems remain insufficiently rigorous for tasks such as structured literature review. This survey categorizes systems, tools, datasets, and evaluation metrics while reviewing applications and challenges across scientific domains. Its analysis finds that literature review remains a significant challenge across nearly all approaches, motivating greater attention to collaboration, reliability, and ethical safeguards.
Problem
Existing Agentic AI systems support scientific workflows but often lack domain-specific focus and compliance-driven rigor, while literature review remains difficult to automate.
Method
The survey categorizes Agentic AI systems and reviews their frameworks, datasets, implementation tools, evaluation metrics, applications, and open challenges.
Results
Literature review exhibits the highest failure rate among several reported research workflow phases and remains a significant challenge across nearly all approaches.
Takeaways & Limitations
Future Agentic AI research should emphasize reliable human-AI collaboration, oversight, and risk evaluation in scientific workflows.
Takeaways & Limitations
Current literature-review frameworks struggle with deep domain knowledge, nuanced understanding, human-AI collaboration, and generalizability across domains.
Abstract
from arXiv · showhide
The integration of Agentic AI into scientific discovery marks a new frontier in research automation. These AI systems, capable of reasoning, planning, and autonomous decision-making, are transforming how scientists perform literature review, generate hypotheses, conduct experiments, and analyze results. This survey provides a comprehensive overview of Agentic AI for scientific discovery, categorizing existing systems and tools, and highlighting recent progress across fields such as chemistry, biology, and materials science. We discuss key evaluation metrics, implementation frameworks, and commonly used datasets to offer a detailed understanding of the current state of the field. Finally, we address critical challenges, such as literature review automation, system reliability, and ethical concerns, while outlining future research directions that emphasize human-AI collaboration and enhanced system calibration.
1 INTRODUCTION
Agentic AI systems are emerging as autonomous tools for automating complex scientific research workflows. This survey reviews these systems and identifies domain-specific rigor as an important gap in existing research automation.
- Agentic AI can independently perform hypothesis generation, literature review, experimental design, and data analysis.These capabilities distinguish agentic systems from traditional AI and support research automation across chemistry, biology, and materials science.
- Existing LLM-driven frameworks automate citation management, document discovery, academic surveys, experimentation, and report writing.Examples include LitSearch, ResearchArena, SciLitLLM, CiteME, ResearchAgent, and Agent Laboratory.
- Many current systems lack the domain-specific focus and compliance-driven rigor needed for structured biomedical evidence synthesis.The passage identifies this limitation in systems designed for general research workflows.
- The survey categorizes Agentic AI systems into autonomous and collaborative frameworks and examines their datasets, implementation tools, and evaluation metrics.It also discusses open challenges to encourage more reliable and impactful scientific contributions.
2 AGENTIC AI: FOUNDATIONS AND KEY CONCEPTS
Agentic AI describes autonomous entities that perceive inputs and take contextually relevant actions through integrated processes such as learning, planning, and decision-making. Single-agent systems suit well-defined tasks, whereas multi-agent systems support collaborative problems requiring multiple runs.
- An AI agent is an autonomous intelligent entity that responds to sensory input with contextually relevant actions.The concept applies across physical, virtual, and mixed-reality environments.
- Agentic AI integrates autonomy, learning, memory, perception, planning, decision-making, and action within interactive systems.
- Single and multi-agent architectures: Single-agent architectures are suited to well-defined problems when user feedback is not required.A single LM-based agent can reason, plan, and execute tools independently across multiple tasks and domains.
- Single and multi-agent architectures: Multi-agent architectures support collaborative problems requiring multiple runs and interactions among specialized agents.Their effectiveness depends on careful interoperability in communication and information sharing.
3 TAXONOMY OF AGENTIC AI FOR SCIENTIFIC DISCOVERY
Agentic AI for scientific discovery spans human-augmented workflows and autonomous systems across domains including chemistry, biology, materials science, and healthcare. Fully autonomous systems can automate end-to-end workflows, while collaborative systems combine AI computation with researcher expertise.
- Agentic AI systems are categorized by autonomy, researcher interaction, and application scope across scientific domains.The survey covers applications from chemistry and biology to materials science and healthcare.
- Fully autonomous systems: Fully autonomous systems operate independently with minimal human intervention across workflows from hypothesis generation to experiment execution.
- Fully autonomous systems: Coscientist plans, designs, and executes chemical experiments using GPT-4.
- Fully autonomous systems: ChemCrow extends GPT-4 with 18 expert-designed tools for organic synthesis, drug discovery, and materials design.
- Human-AI collaborative systems: Human-AI collaborative systems combine AI computational power with researchers’ creativity and expertise.Virtual Lab organizes team meetings and individual tasks for interdisciplinary research, including nanobody binder design for SARS-CoV-2.
4 AGENTIC AI FOR LITERATURE REVIEW
Literature review is foundational to scientific discovery but has become difficult as publications grow rapidly. Agentic AI can automate retrieval, extraction, and synthesis, yet current systems still struggle with domain knowledge, nuanced understanding, collaboration, and generalizability.
- Literature review helps researchers identify trends, evaluate methods, recognize knowledge gaps, frame questions, and support reproducibility.Its importance spans chemistry, biology, materials science, healthcare, and artificial intelligence.
- The growth of scientific publications has made traditional manual literature reviews increasingly challenging.Researchers therefore use autonomous agents to navigate large literature datasets and extract relevant information.
- Agentic AI automates information retrieval, extraction, and synthesis for literature review.Related technologies also support automatic extraction, trend analysis, and predictive modeling.
- Existing frameworks: SciLitLLM combines continual pre-training and supervised fine-tuning to improve domain-specific literature understanding and instruction following.It improves document classification, summarization, and question answering but depends heavily on high-quality training data.
- Open challenges: Literature-review automation remains limited by deep domain knowledge, nuanced understanding, insufficient human-AI collaboration, and restricted cross-domain generalizability.Agent Laboratory showed a significant performance drop during literature review, while many frameworks prioritize fully autonomous workflows or target specific domains.
5 AGENTIC AI FOR SCIENTIFIC DISCOVERY
Agentic AI is being applied across the scientific research lifecycle, from ideation and experiment execution to data analysis, writing, and dissemination. Applications span chemistry, biology, materials science, and general science, with systems such as Coscientist demonstrating increasingly integrated workflows.
- Research lifecycle: Agentic AI systems automate research stages including ideation, experiment design and execution, data analysis, and paper dissemination.These systems use literature analysis, planning, optimization, robotic automation, large-scale data processing, and manuscript generation across the workflow.
- Chemistry: Chemistry applications include molecular discovery, reaction prediction, synthesis planning, laboratory automation, and computational chemistry.The survey describes agents that accelerate compound identification, optimize synthetic routes, execute experiments robotically, and run molecular simulations.
- Chemistry: Coscientist plans, designs, and executes chemical experiments through web search, documentation analysis, code execution, and robotic automation.It demonstrated multi-step problem-solving by designing and optimizing a palladium-catalyzed cross-coupling reaction.
- Biology: Biology agents automate data analysis, hypothesis generation, and experimental planning using genomic, protein, and biomedical information.The survey links these capabilities to genetics, drug discovery, and synthetic biology, while BIA focuses on scRNA-seq workflows and report generation.
- Additional domains: Agentic AI research also covers materials science, general science, and machine learning, alongside frameworks summarized in Figure 2.The cited work indicates that the field extends beyond chemistry and biology into multiple scientific domains.
6 IMPLEMENTATION TOOLS, DATASETS AND METRICS
Scientific agent development depends on computational frameworks, curated datasets, and task-specific evaluation metrics. The survey organizes these resources around agent construction, scientific benchmarks, and multidimensional assessment.
- Implementation tools: Agentic AI systems combine foundational models, computational frameworks, and domain-specific tools to execute scientific tasks.These resources support both single-agent and multi-agent architectures.
- Implementation tools: AutoGen manages customizable, conversable multi-agent systems, while MetaGPT assigns specialized roles using an assembly-line workflow.Letta supports persistent agent services and explicitly incorporates cognitive architecture principles.
- Datasets: Scientific datasets evaluate agents’ reasoning, planning, collaboration, hypothesis generation, literature analysis, and experimental planning.LAB-Bench and MoleculeNet provide examples for benchmarking biological and chemical data understanding.
- Metrics: Evaluation metrics vary by task, including accuracy, task completion, response coherence, precision, recall, prediction error, explainability, and human evaluation.The appropriate metric depends on the scientific domain and application.
7 CHALLENGES AND OPEN PROBLEMS
Agentic AI for scientific discovery faces reliability, trustworthiness, ethical, and operational challenges. The survey emphasizes realistic benchmarking, human oversight, and safeguards against bias, unsafe behavior, and cascading system failures.
- Trustworthiness: Trustworthy evaluation should reflect real-world conditions while jointly considering accuracy, cost, speed, throughput, reliability, and recovery from failure.The survey also emphasizes predictability, explainability, and safety rather than overly complex or costly designs.
- Ethical and practical considerations: Ethical risks include bias, privacy, accountability, compliance, transparency, fairness, and misleading or fabricated LLM outputs.These concerns are especially consequential when agents are deployed in critical domains such as healthcare.
- Ethical and practical considerations: Multi-agent and tool-using systems require oversight because one unethical or misaligned agent can compromise the integrity of the entire system.Proposed responses include human-in-the-loop architectures and mechanisms for evaluating and mitigating risks during training and deployment.
- Potential risks: Unreliable or biased data, inadequate human oversight, goal misalignment, coordination failures, protocol deviations, and missed safety measures can produce harmful scientific outcomes.The survey highlights compounded errors, irrelevant experiments, irreproducible findings, hazardous actions, and the need to define the agents’ blast radius.
- Potential risks: The field must monitor predictability and control as autonomy increases, particularly when agents are connected to robotic laboratories.This boundary follows from the possibility that unintended actions may be difficult to detect or correct in real time.
8 CONCLUSION AND FUTURE DIRECTIONS
The survey identifies literature review as a major unresolved challenge across agentic AI approaches, while organizing the field’s systems, frameworks, datasets, benchmarks, and open problems. It highlights calibration as a future direction for improving output accuracy and reliability.
- The survey systematically reviews agentic AI approaches for scientific discovery across chemistry, biology, materials science, and related domains.
- It organizes the field through a taxonomy of functional frameworks, literature-review approaches, datasets, and benchmarks.
- Literature review remains a significant challenge across nearly all approaches, particularly for research idea generation and scientific discovery.
- The literature-review phase exhibited the highest failure rate among data preparation, experimentation, report writing, and research report generation.
- Future work should integrate calibration techniques so agents’ confidence aligns more closely with the actual correctness of their predictions.