Source-linked AI summary
Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
Steffen Eger, Yong Cao, Jennifer D'Souza, Andreas Geiger, Christian Greisinger, Stephanie Gross, Yufang Hou, Brigitte Krenn, Anne Lauscher, Yizhi Li, Chenghua Lin, Nafise Sadat Moosavi, Wei Zhao, Tristan Miller
TL;DR
AI systems are increasingly being applied across the scientific research lifecycle, but the evidence and evaluation practices remain fragmented. This survey synthesizes representative methods across five research tasks, finding that AI can support selected components of scientific work while still requiring human oversight and careful safeguards.
Problem
AI-assisted science spans multiple research tasks, but existing evidence, evaluation practices, and ethical guidance need a cross-task synthesis.
Method
The survey reviews representative datasets, methods, results, evaluation strategies, limitations, and ethical concerns across five aspects of the research cycle.
Results
AI systems can meaningfully support certain components of scientific research, but AI4Science should presently complement rather than replace human knowledge creation.
Takeaways & Limitations
AI-assisted tools may enhance the accessibility and reliability of research when integrated as complementary systems rather than replacements for human scientific work.
Takeaways & Limitations
Factual consistency, truthfulness, and bibliographic citations require human oversight, while reliable task-specific evaluation methods remain limited.
Abstract
from arXiv · showhide
With the advent of large multimodal language models, science is now at a threshold of an AI-based technological transformation. An emerging ecosystem of models and tools aims to support researchers throughout the scientific lifecycle, including (1) searching for relevant literature, (2) generating research ideas and conducting experiments, (3) producing text-based content, (4) creating multimodal artifacts such as figures and diagrams, and (5) evaluating scientific work, as in peer review. In this survey, we provide a curated overview of literature representative of the core techniques, evaluation practices, and emerging trends in AI-assisted scientific discovery. Across the five tasks outlined above, we discuss datasets, methods, results, evaluation strategies, limitations, and ethical concerns, including risks to research integrity through the misuse of generative models. We aim for this survey to serve both as an accessible, structured orientation for newcomers to the field, as well as a catalyst for new AI-based initiatives and their integration into future ``AI4Science'' systems.
1 Introduction
AI systems are increasingly positioned to support the full scientific research cycle, from literature search and experimentation to content generation and evaluation. This survey reviews representative approaches across these tasks while emphasizing their limitations and ethical risks.
- 148,000 papers across 22 non-CS disciplines show rapidly increasing citations of large language models between 2018 and 2024.
- Researchers widely expect AI use to become mainstream within two years, although current use is often limited to writing assistance.
- AI tools now support multiple stages of research, including literature search, experimentation, multimodal content generation, and automated peer review.
- The survey organizes representative work around five tasks: literature search, experimentation and idea generation, textual generation, multimodal generation, and peer review.
- Current systems face hallucination, bias, limited reasoning, environmental costs, weak evaluation, fake science, plagiarism, and reduced human oversight.
- The survey addresses ethical concerns alongside each research task and in a dedicated discussion of broader integrity risks.
2 Survey Scope and Methodology
The survey uses a broad, workflow-centric and narrative approach to synthesize representative AI-assisted science research across domains. Its selection strategy prioritizes methodological clarity, comparability, and relevance while acknowledging that coverage is not exhaustive.
- The survey provides a broad, workflow-centric overview intended to orient researchers across the full research lifecycle.
- Its narrative methodology is designed for synthesizing heterogeneous, rapidly evolving work and drawing connections across domains and methods.
- The survey is not exhaustive and does not claim to capture every recent publication.
- Representative works were selected by domain-expert teams using seed papers, citation analysis, keyword searches, and criteria including relevance, methodological clarity, venue reputation, and impact.
- The survey favors approaches with core techniques, well-defined evaluation protocols, or influence on subsequent work, supporting comparison across tasks.
- A common subsection structure highlights evaluation dimensions, methodological trade-offs, and connections between tasks while reducing single-viewpoint selection bias.
3 AI Support for Individual Topics and Tasks
The survey presents AI support across literature search, idea and hypothesis generation, experimentation, scientific text generation, and evaluation. Results show benefits from structured retrieval, iterative and collaborative generation, while factuality, generalization, feasibility, and reproducibility remain important constraints.
- Literature Search: Retrieval systems support complex scientific questions, bibliometric analysis, field mapping, similarity search, and document-grounded answers.Retrieval-augmented responses are grounded in provided documents, improving factual reliability and interpretability while helping mitigate hallucinations.
- Literature Search: Search tools remain vulnerable to incomplete, outdated, or biased sources, scalability constraints, proprietary backends, and the cold-start problem.These limitations affect retrieval accuracy, diversity of perspectives, reproducibility, and recommendations for new papers or users.
- Idea & Hypothesis Generation: Iterative refinement, multi-agent teamwork, and structured knowledge can improve idea and hypothesis generation, including novelty, testability, and research-direction discovery.However, fine-tuning gains may remain domain-limited, and unseen-domain fine-tuning can harm hypothesis quality, particularly novelty.
- Idea & Hypothesis Generation: LLM-generated ideas can be more novel but slightly less feasible, while expert evaluations have found generated ideas less novel, exciting, or effective than human ideas.The evidence indicates a persistent trade-off between novelty and feasibility and uncertainty about whether theoretical ideas lead to scientific discovery.
- Experimentation & Evaluation: Across experimentation and evaluation, LLMs face hallucinated results or references, weak multimodal alignment, insufficient critical analysis, and challenges with precise reasoning and tool use.The survey therefore emphasizes human oversight for factual consistency, truthfulness, and bibliographic citations, while peer-review evaluation remains domain-limited.
4 Ethical Concerns
The survey frames ethical concerns around truthfulness, bias, privacy, academic integrity, and responsible disclosure in AI-assisted scientific work. It emphasizes that AI-generated research content still requires human oversight and accountability.
- Ethics research on generative AI covers hallucination, alignment, harmful content, copyright, private-data leakage, creativity, safety, fairness, privacy, and robustness.
- Truthfulness is a central challenge because scientific use requires accurate representation of information, facts, and results.
- TrustLLM evaluations generally find proprietary LLMs outperform open-source models on trustworthiness, with Llama2 as a notable exception.
- Models often struggle to produce truthful responses from internal knowledge alone, while external knowledge improves performance and correlates positively with downstream functional effectiveness.
- Scientific authors remain responsible for manuscript integrity, accuracy, fairness, privacy, copyright, and transparent disclosure of AI use, while AI cannot be listed as an author.
- Editorial guidance described in the survey prohibits AI use in reviewing and limits AI-generated images or multimedia to cases where they are specifically allowed.
5 Conclusion
The survey reviews AI support across five parts of the scientific research cycle, while emphasizing that current systems remain limited, uneven, and dependent on human oversight. It presents AI4Science as complementary to human expertise rather than a replacement.
- The survey covers search, experimentation and idea generation, text production, multimodal content production, and peer review.
- For each topic, it discusses datasets, methods, results, evaluation strategies, limitations, ethical concerns, and future research directions.
- AI systems can meaningfully support some scientific-workflow components, but their capabilities remain limited and uneven.
- Many methods rely on narrow benchmarks, struggle with generalization, or require substantial human oversight to avoid errors, bias, and misinterpretation.
- The survey concludes that AI4Science should currently be viewed as complementary tools rather than a transformative replacement for human expertise.
- The authors hope AI4Science initiatives will support faster, more efficient, inclusive, accessible, and reliable scientific discovery while maintaining ethical standards.
A Historical Context and Background
Scientific research proceeds through an iterative cycle from questions and literature review to hypotheses, experiments, analysis, reporting, and further inquiry. Modern publication growth makes maintaining familiarity with existing research increasingly difficult.
- The scientific cycle includes forming a question, studying existing literature or data, formulating a falsifiable hypothesis, testing it, analyzing results, and reporting findings.
- Scientific publications have been doubling every 17 years, making exhaustive reading of the literature unworkable.
- Hypotheses commonly emerge through successive cycles of generation, evaluation, and modification or rejection rather than induction alone.
- Experiments and analysis aim to establish causal relationships between variables relevant to a scientific hypothesis.
- Experimental design includes randomization, replication, blocking, and advance determination of statistical analysis, subject to resource constraints.
- Reporting disseminates findings through articles, books, and presentations, while peer review addresses evaluation, reliability, objectivity, and bias.
B.1 Literature Search, Summarization, and Comparison
Scientific search systems range from conventional indexes to AI-enhanced platforms that organize literature, code, datasets, and research knowledge. These tools can accelerate discovery but raise concerns about ranking bias, transparency, accountability, and equity.
- Traditional academic search engines provide broad scholarly access but offer limited filtering and relatively basic relevance ranking compared with AI-enhanced tools.
- Scientific search platforms aggregate documents, provide citation analysis, and support exploration of influential works and emerging trends.
- Code- and dataset-focused platforms connect papers with implementations and datasets, supporting reproducibility, comparison, and practical application.
- Dataset discovery tools help researchers identify resources for specific problems, fostering collaboration and accelerating experimentation cycles.
- AI-assisted search can automate discovery and uncover patterns, but it also introduces risks involving transparency, accountability, equity, bias, and reinforcement of the Matthew effect.
- The survey favors human-centric search in which researchers use advanced tools while retaining responsibility for conducting research and summarizing results.
- Content-based recommendations and literature-gap detection are presented as ways to reduce popularity-driven bias and encourage attention to underpopulated topics.
B.2 AI-Driven Scientific Discovery: Ideation, Hypothesis Generation, and Experimentation
AI-assisted scientific discovery combines methods for generating hypotheses and ideas with automated experimentation. These approaches draw on diverse information sources, iterative refinement, multi-agent collaboration, and computational resources, while raising concerns about bias, transparency, safety, and oversight.
- Methods: Figure 5 presents a broad overview of methods used for hypothesis generation, idea generation, and automated experimentation.
- Methods: Hypothesis and idea generation use iterative refinement, alignment strategies, and multi-agent collaboration, while automated experimentation uses tree search, multi-agent workflows, and refinement.Initial hypotheses can be validated against knowledge bases, and long contexts can be summarized and integrated during refinement.
- Methods: Hypothesis and idea generation leverage scientific literature, web data, and datasets, whereas automated experimentation requires computational models, simulations, and raw data.
- Ethical Concerns: LLM-based ideation may reinforce established research paradigms by favoring popular paths and neglecting underrepresented directions.This can marginalize unconventional ideas and reduce the diversity of scientific thinking.
- Ethical Concerns: LLM-generated hypotheses may lack transparency about their validity and underlying assumptions, making scientific soundness and accountability difficult to assess.Correlation-based hypotheses may be proposed without clearly revealing their assumptions, potentially leading to flawed experiments.
- Ethical Concerns: Ideation and hypothesis-generation systems are not safe by design and may be jailbroken to produce harmful ideas, toxic molecular designs, or unsafe procedures.Examples include improper equipment use, unsafe chemical handling, and failure to recognize experimental hazards.
B.3 Text-based Content Generation
Text-based scientific content generation covers titles, abstracts, related work, and bibliographies through extractive, abstractive, retrieval-based, and parametric methods. However, AI-generated text raises authorship, plagiarism, attribution, and detection concerns.
- Methods: Academic paper content generation includes title, abstract, related-work, citation, and bibliography generation.Figure 6 illustrates the overall content-generation process for academic papers.
- Methods: Title generation maps abstracts, content, or future work to titles, while abstract generation maps titles or keywords to abstracts.
- Methods: Related-work generation uses extractive sentence reordering or abstractive rewriting across multiple papers.
- Methods: Bibliography generation is divided into non-parametric retrieval from external sources and parametric generation from preexisting model knowledge.Non-parametric approaches may retrieve references before generation or check for citations afterward.
- Ethical Concerns: AI-generated text is difficult to distinguish from human-written text, and automatic detectors can be fooled by paraphrasing.
- Ethical Concerns: ChatGPT-generated texts can pass automated plagiarism detectors, complicating authorship and plagiarism assessment.
B.4 Multimodal Content Generation and Understanding
The survey identifies datasets for multimodal content generation and understanding as a component of scientific multimodal research.
- Data: Table 3 provides an overview of datasets for multimodal content generation and understanding.
B.4.1 Data.
Scientific multimodal datasets support table understanding, table generation, literature-review table construction, numerical reasoning, and scientific figure generation. These resources increasingly ground outputs in surrounding text and evaluate numerical interpretation across modalities.
- Scientific Table Understanding: Scientific table understanding commonly involves generating text from tables, including descriptions grounded in surrounding scientific context.SciXGen draws from over 200K scientific papers and covers descriptions of tables, figures, and algorithms.
- Scientific Table Understanding: Numerical table benchmarks include expert-annotated tables paired with paper passages describing their findings.Each dataset contains 1.3K expert-annotated tables.
- Scientific Table Understanding: Some datasets pair numerical ML-performance tables with expert-written explanations of precision, recall, and accuracy.
- Scientific Table Understanding: Hierarchical-table datasets introduce numerical reasoning tasks for complex tables commonly found in statistical reports.
- Scientific Table Understanding: Recent work shows that scientific table performance depends strongly on the table source or modality, such as PDF-rendered images versus LATEX/HTML tables.Dedicated multimodal benchmarks target numerical reasoning.
- Scientific Table Generation: Scientific table generation converts unstructured text into structured tables, including literature-review tables whose rows represent papers and columns capture methods, datasets, and results.ArXivDIGESTables uses captions and in-text context as additional grounding information.
- Scientific Figure Generation: Scientific figure-generation tools accept sketches, screenshots, or text, generate TikZ code, and render high-quality vector graphics.
B.4.3 Ethical Concerns.
Multimodal scientific-content tools face data limitations that can increase hallucination risks and produce incorrect figures, especially when users overlook or misuse those limitations.
- Data limitations: AutomaTikZ and its extensions contain only several hundred thousand text–code pairs, whereas general-purpose image-generation datasets are often orders of magnitude larger.This creates a substantial scale difference between specialized scientific-figure data and general image-generation resources.
- Data limitations: Misalignment between image captions and corresponding images or code increases the risk of hallucinations.The cited limitation concerns correspondence quality between textual descriptions and visual or code outputs.
- Practical risks: These tools can produce incorrect scientific figures when users overlook, ignore, or maliciously abuse their limitations.The risk is tied both to model limitations and to how users interact with the systems.
B.5 Peer Review
AI-assisted peer review and scientific evaluation span rigor assessment, claim verification, and presentation analysis, but remain constrained by incomplete evidence and ethical risks. The survey emphasizes contextualized verification and human oversight rather than relying on narrow or fully automated systems.
- Evaluation scope: Automated scientific evaluation addresses rigor, claim validity, evidence retrieval, and how researchers frame findings.The surveyed approaches include rigor criteria, scientific fact verification, contextual evidence retrieval, and detection of overstatement or understatement.
- Rigor assessment: Rigor-assessment frameworks combine keyword extraction, definition generation, and salient-criterion identification, with domain-agnostic designs enabling adaptation across fields.These components support computational analysis of rigor while allowing flexible use across domains.
- Claim verification: SciFact-Open provides claims and supporting evidence from abstracts, but abstract-only evidence can be inaccurate or misleading and should be corroborated with the paper body.The survey contrasts abstract-based verification with the need for evidence from the main text.
- Claim verification: A lab-note claim dataset links claims to figures, tables, and methodological details to support context-based verification, but primarily evaluates factual correctness rather than overstatement.Its annotations and retrieval tasks provide richer context, while the survey identifies a remaining limitation in evaluating presentation of findings.
- Ethical concerns: AI-supported peer review raises ethical concerns including unfair affiliation bias, diminished reviewer engagement, and reduced critical thinking.The survey reports affiliation bias in abstract review and highlights concerns about reviewers’ engagement with and critical assessment of manuscripts.
- Human oversight: Systems targeting particular aspects while preserving human scientific autonomy may be preferable to end-to-end reviewing systems.The survey frames human autonomy as a design consideration for safer collaborative peer-review support.
- Human oversight: The survey calls for ethical standards that balance AI capabilities with human expertise in scientific review.This conclusion follows concerns about bias, reviewer engagement, and critical thinking.