Source-linked AI summary

From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery

Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, Yangqiu Song

arXiv:2505.13259v3cs.CL

TL;DR

Scientific discovery is increasingly incorporating LLMs whose roles extend beyond task automation toward more autonomous research processes. This survey organizes the shift across the scientific method with a three-level taxonomy and identifies challenges and future directions for autonomous discovery.

  • Problem

    Existing reviews often focus on specific scientific domains or static LLM capabilities, leaving increasing autonomy and evolving roles across the scientific method insufficiently structured.

  • Method

    The survey systematically analyzes LLM applications across six scientific-method stages using the taxonomy LLM as Tool, Analyst, and Scientist.

  • Results

    LLMs are progressing from discrete, task-oriented functions toward sophisticated multi-stage agentic workflows, with emerging systems navigating nearly all scientific-method stages.

  • Takeaways & Limitations

    The survey identifies autonomous discovery cycles, robotic automation, self-improvement, transparency, interpretability, and ethical governance as pivotal future directions.

  • Takeaways & Limitations

    The survey excludes exhaustive coverage of general-purpose scientific LLMs, domain-specific scientific reasoning, and orthogonal benchmarks for fundamental LLM capabilities.

Abstract

from arXiv · show

Large Language Models (LLMs) are catalyzing a paradigm shift in scientific discovery, evolving from task-specific automation tools into increasingly autonomous agents and fundamentally redefining research processes and human-AI collaboration. This survey systematically charts this burgeoning field, placing a central focus on the changing roles and escalating capabilities of LLMs in science. Through the lens of the scientific method, we introduce a foundational three-level taxonomy-Tool, Analyst, and Scientist-to delineate their escalating autonomy and evolving responsibilities within the research lifecycle. We further identify pivotal challenges and future research trajectories such as robotic automation, self-improvement, and ethical governance. Overall, this survey provides a conceptual architecture and strategic foresight to navigate and shape the future of AI-driven scientific discovery, fostering both rapid innovation and responsible advancement. Github Repository: https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery.

1 Introduction

The survey addresses the need for a framework that tracks LLM applications across the scientific method and their increasing autonomy. It introduces a three-level taxonomy and identifies future challenges for autonomous scientific discovery.

  • Existing reviews often provide discipline-specific coverage or static capability snapshots, overlooking increasing LLM autonomy across the scientific method.
  • The survey organizes analysis around six scientific-method stages, from observation and problem definition through iteration and refinement.
  • LLMs are progressing from discrete, single-stage tasks toward sophisticated multi-stage agentic workflows.
  • The taxonomy distinguishes LLM as Tool, Analyst, and Scientist according to escalating autonomy and independence in research.
  • The survey highlights autonomous discovery cycles, robotic experimentation, self-improvement, transparency, interpretability, and ethical governance as key future directions.
  • The survey excludes general-purpose scientific LLMs and domain-specific knowledge acquisition or reasoning because existing surveys already cover those areas.

2 Three Levels of Autonomy

The survey defines three autonomy levels for LLM-based scientific discovery, ranging from supervised task execution to largely independent orchestration of research stages.

  • LLM as Tool: Level 1, LLM as Tool, covers specific, well-defined tasks within one scientific-method stage under direct human supervision.
  • LLM as Analyst: Level 2, LLM as Analyst, supports complex information processing, data modeling, and analytical reasoning with reduced human intervention.
  • Figure 2 consolidates research works within the survey’s focused scope across all three autonomy levels.
  • LLM as Analyst: Analyst systems can independently manage task sequences, while researchers define goals, provide data, and critically evaluate generated insights.
  • LLM as Scientist: Level 3, LLM as Scientist, involves active agents that orchestrate multiple discovery stages with considerable independence.
  • LLM as Scientist: Scientist-level systems may formulate hypotheses, plan and execute experiments, analyze data, draw preliminary conclusions, and propose subsequent research questions.

3 Level 1. LLM as Tool (Table A1)

Level 1 research treats LLMs as tools supporting discrete scientific-method activities, including literature work, idea generation, experimentation, data analysis, review, and hypothesis validation.

  • Literature Review and Information Aggregation: Literature-review tools automate scientific search, retrieval, summarization, and aggregation into structured tables.
  • Idea Generation and Hypothesis Formulation: LLMs support automated generation of research ideas and testable hypotheses from literature summaries and open-ended prompts.
  • Experiment Planning and Execution: At Level 1, LLMs assist experiment planning and execution, with execution research concentrated on code generation for computational environments.
  • Data Analysis: Level 1 applications process tabular and chart data for organization, presentation, analysis, and question answering.
  • Conclusion and Hypothesis Validation: LLMs can review papers, identify inserted errors, and support multi-agent or reinforcement-learning approaches intended to improve review robustness.
  • Conclusion and Hypothesis Validation: Hypothesis-validation systems generate verification code, evaluate replication, predict empirical outcomes, and use multi-agent workflows in physics research.
  • Hypothesis Refinement: Iterative refinement of research hypotheses has received comparatively less attention than other Level 1 activities.

4 Level 2: LLM as Analyst (Table A2)

Level 2 research examines LLMs as passive agents that handle more complex analytical sequences across machine learning, data analysis, function discovery, and scientific domains.

  • Level 2 research covers works categorized by task nature and scientific domain.
  • Automated Machine Learning: In automated machine learning, benchmarks evaluate LLMs’ ability to design and execute experiments, with performance often contingent on task familiarity.
  • Automated Machine Learning: Agentic frameworks address benchmark challenges through iterative refinement, machine-generated ideas, expert collaboration, and end-to-end modeling systems.
  • Automated Data-Driven Analysis: Data-analysis benchmarks indicate that most LLMs struggle with complex analytics tasks even when operating within agent frameworks.
  • Function Discovery: Function-discovery systems use LLM domain knowledge, clustered memory feedback, and dual reasoning to identify equations from observational data.
  • Domain Applications: Scientific-domain research includes chemistry and social-science causal discovery, biomedical research, drug discovery, chemistry experimentation, and biochemistry discovery.
  • Cross-Domain Benchmarks: Broad benchmarks evaluate diverse scientific-discovery tasks in virtual environments and other agent-based research scenarios.

5 Level 3. LLM as Scientist (Table A3)

Level 3 systems pursue increasingly autonomous scientific workflows, from broad human inputs through iterative refinement and research outputs. Their feedback loops combine automated critique with human guidance for strategic redirection.

  • End-to-End Workflow: These systems use agent-based frameworks to span workflows from literature review through iterative refinement and draft research outputs.
  • Research Genesis: Level 3 systems can begin from reference papers or broad research domains rather than narrowly specified objectives.
  • Iterative Refinement: Iterative refinement can trigger fundamental reassessments of hypotheses, designs, or the broader research trajectory.
  • Iterative Refinement: The AI Scientist uses automated reviewers, LLM evaluators, and vision-language models to create an internal feedback loop.
  • Iterative Refinement: Other systems integrate human expertise for macro-level guidance, enabling strategic redirection when results challenge the research premise.

6 Challenges and Future Directions

The survey identifies future directions needed to extend autonomous scientific discovery across research cycles, physical experimentation, adaptive learning, interpretability, and governance. These directions address both capability gaps and the conditions for reliable deployment.

  • Fully-Autonomous Research Cycle: Future systems should pursue autonomous research cycles that identify follow-up questions and direct efforts beyond a single predefined inquiry.
  • Robotic Automation: Integrating LLMs with robots could enable physical laboratory experimentation and broaden end-to-end discovery in chemistry and materials science.
  • Transparency and Interpretability: Reliable and reproducible discovery requires interpretable systems whose internal reasoning supports verifiable claims and scientifically justifiable conclusions.
  • Continuous Self-Improvement: Continual learning and online reinforcement learning are proposed to help scientific agents assimilate outcomes and adapt strategies over their operational lifetimes.
  • Ethics and Societal Alignment: As systems gain independent reasoning and action, adaptive governance and embedded ethical constraints are needed to address misuse, bias, and human-control risks.

Limitations

The survey focuses on how LLM autonomy changes across the scientific method, rather than comprehensively reviewing every capability or domain-specific scientific LLM. Its exclusions are deliberate and preserve this autonomy-centered scope.

  • The review categorizes scientific-discovery systems as LLM as Tool, LLM as Analyst, or LLM as Scientist across the scientific method.
  • It excludes exhaustive coverage of general-purpose scientific LLMs for domain-specific reasoning or application.
  • The survey also does not deeply examine orthogonal benchmarks or methodologies for general abilities such as planning, code generation, and agentic decision-making.

Ethics Statement

The paper presents a comprehensive survey of LLMs in scientific discovery and reports that its reviewed works are cited and publicly accessible or research-licensed. It did not conduct additional dataset curation or human annotation.

  • The survey examines LLMs’ transformation from task automation tools to autonomous agents in scientific discovery.
  • The reviewed research works are properly cited and, to the authors’ knowledge, publicly accessible or available under research-use licenses.
  • The paper did not conduct additional dataset curation or human annotation work.

A Summary Tables of LLMs in Scientific Discovery

Tables A1–A3 compare and classify research works in LLM-based scientific discovery across the survey’s three taxonomy levels.

  • Table A1 compares and classifies Level 1 research works in LLM-based scientific discovery.
  • Table A2 compares and classifies Level 2 research works in LLM-based scientific discovery.
  • Table A3 compares and classifies Level 3 research works in LLM-based scientific discovery.
Loading 2505.13259v3…