Source-linked AI summary
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
Patrick Emami, Sameera Horawalavithana, Truc Nguyen, Gihan Panapitiya, Bruno Jacob, Siddhisanket Raskar, Saumya Sinha, Jared D. Willard, Andrew Glaws, Nithin Somasekharan, Ling Yue, Brian Lu, Shaowu Pan, Jason Eisner
TL;DR
Research on AI Scientists has largely emphasized autonomous capability rather than the human-agent dynamics of scientific teamwork. This paper reviews AI Scientists through a human-agent systems lens and analyzes case studies, finding that human expertise and agents can augment each other’s outputs while motivating mathematical frameworks for human-AI synergy.
Problem
Research on AI Scientists underexplores the human-agent pair and broader social dynamics of scientific teamwork.
Method
The paper reviews AI Scientist literature through the human-agent systems lens and analyzes real-world scientist-agent case studies.
Results
Case studies qualitatively illustrate humans and agents augmenting each other’s outputs and improving scientific outcomes.
Takeaways & Limitations
Studying AI Scientists as human-agent systems offers a route toward safer and more effective human-agent co-discovery in science.
Takeaways & Limitations
The practical value of human-agent systems may be offset by expensive, difficult-to-reproduce evaluations and unmodeled measurement costs.
Abstract
from arXiv · showhide
Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of "AI Scientists". We argue that this overlooks the social aspects of scientific teamwork, and that studying AI Scientists as human-agent systems (HAS)--where the unit of analysis is the human-agent pair--is both underexplored and undervalued. We establish these points through literature and empirical analysis, and highlight recent incidences and studies which show that deploying agents in science without accounting for human-agent dynamics introduces near-term risks, including reduced diversity of scientific inquiry. Through analysis of real-world case studies, we show that scientists and agents can augment each other's capabilities. We call for new research that adopts the HAS lens to develop mathematical frameworks for understanding and fostering human-AI synergy in scientific discovery.
1 Introduction
Scientific teamwork is increasingly incorporating LLM-based agents as collaborators, but existing research largely emphasizes autonomous capabilities rather than sustained human-agent collaboration. The section argues for studying AI Scientists as human-agent systems to understand risks and enable scientific outcomes that humans or agents alone may not achieve.
- Scientific collaboration combines diverse expertise to address complex questions across teams ranging from small laboratories to large institutions.
- LLM-based agents are shifting from tools toward collaborators in scientific processes amid growing scientific teamwork and automation.
- Existing systems often support goal setting and work verification but less often sustain the iterative, joint collaboration characteristic of human teamwork.
- The proposed human-agent systems lens treats humans and agents as the unit of study within an interface-mediated environment.An HAS includes agents, human participants, an interaction interface, and the surrounding environment.
- AI-augmented research produced three times as many papers with a 5% reduction in explored-topic diversity, motivating HAS research on deployment risks and human-agent synergy.The section also points to case studies where scientists and agents augmented each other’s outputs and improved scientific outcomes.
2 AI Scientists as HAS: A Review and Preliminary Experiment
The review applies a human-agent systems lens to classify AI Scientists by how human feedback enters discovery. It finds that most systems keep humans in supervisory roles, while flexible asynchronous interaction remains uncommon; a preliminary benchmark experiment tests whether agents can solicit useful feedback.
- Literature review: The review classifies AI Scientists by feedback channel: artifact-level, discovery phase-level, research planning, and asynchronous steering and interruption.These categories organize how human feedback enters the scientific discovery loop.
- Feedback mechanisms: Human feedback in sequential discovery systems is mainly evaluative and corrective, ranging from coarse review gates to fine-grained editing of phase-level memory.Multi-agent systems expose phase-specific roles and review gates, while Denario enables scientists to edit intermediate state before execution resumes.
- Feedback mechanisms: Phase-specific feedback addresses failure modes such as vague hypotheses through textual critique, fine-grained output feedback, ranked idea selection, and hypothesis tournaments.The reviewed systems vary feedback granularity from high-level comments on ideas to detailed intervention on complex scientific outputs.
- Literature review: Most AI Scientists assign humans supervisory roles in scoping, auditing, and approval, treating discovery primarily as an optimization problem weakly connected to human-led research.Artifact-level systems are most numerous, discovery-phase systems form the next tier, planning-centric systems are smaller but growing, and only a handful support flexible asynchronous interaction.
- Preliminary experiment: A preliminary experiment on 10 End-to-End Discovery tasks compares an ask-user tool with increasing feedback guidance to assess agents’ ability to request fine-grained corrective help.Questions receive an automatic continuation prompt under the AstaBench default.
3 Studying AI Scientists as HAS can assist near-term risk mitigation
Studying AI Scientists as human-agent systems can mitigate near-term risks by keeping humans engaged in discovery, oversight, governance, and skill formation. This lens addresses hallucinations, bias, security threats, overloaded review, and unclear accountability.
- Epistemic risks: 51 accepted NeurIPS 2025 papers contained 100 confirmed hallucinated citations, illustrating the need for human oversight of AI Scientist outputs.LLMs and agents can also fabricate data or results on demanding tasks.
- Epistemic risks: End-of-process review is inadequate; asynchronous human interaction, trace-log and code audits, and multi-granular oversight can surface hallucinations and biases earlier.These systems must also account for cognitive load created by AI’s speed and scale.
- Institutional risks: NeurIPS and ICLR submissions increased 10.4× from 2014 to 2024 without commensurate reviewer growth, while AI Scientists expand scientific software and data attack surfaces.Backdoors can induce malicious code, and poisoning crowd-sourced knowledge bases can enable such attacks.
- Institutional risks: Automated governance—including red teaming, adversarial evaluations, monitoring, and provenance tracking—can address misuse and security as AI Scientists are deployed.The HAS lens frames these measures as part of maintaining institutional health.
- Institutional risks: HAS research can help preserve deliberate practice and the skill formation through which scientists develop competencies, because interaction with AI systems affects learning.The paper identifies how scientists interact with AI as central to de-skilling concerns.
- Accountability risks: When AI Scientists support harmful or sensitive research, accountability is unclear amid evolving legal liability and concerns about unintentional plagiarism.Determining originality can be difficult even when generated text is not verbatim.
4 The potential of human-AI synergy in science: case studies
Published case studies show that human-agent pairs can mutually augment scientific work: agents accelerate implementation and drafting, while scientists provide conceptual direction, domain judgment, and epistemic verification. These collaborations can compress complex workflows, but they depend on expert oversight and may homogenize scientific writing or weaken skill development.
- Cross-case pattern: Across the case studies, scientists contributed conceptual framing, domain judgment, and epistemic verification, while agents compressed implementation and also contributed conceptually.The cases were drawn from published accounts of working scientists collaborating with agents on concrete scientific tasks.
- Complexity-theory case: Gemini generated a roadmap, manuscript, and expanded proofs over eight prompt turns, reducing the overhead of turning an informal theorem result into a publishable write-up.The agent drafted in the human author’s writing style and iteratively developed proof details and exposition.
- Complexity-theory case: Human expertise corrected missing context and an erroneous lemma assumption, enabling an expert-plus-junior-collaborator workflow that accelerated drafting and supported a minor-result publication.The human set research direction and quality standards, while the agent accelerated drafting and execution.
- Complexity-theory case: The complexity-theory workflow may homogenize scientific writing and risk impeding early-career scientists’ development of writing and reasoning skills through deliberate practice.The risk arises if this interaction mode is widely adopted, especially by researchers still developing core skills.
- Thermonuclear-burn case: GPT-5 translated a human concept into a PDE-based computational workflow, while the scientist tuned physical parameters, debugged numerics, and rejected noisy or oversimplified outputs.The agent supported model construction, discretization, simulation, experiments, and optimization; the human supplied physically meaningful parameterization and epistemic checks.
- Thermonuclear-burn case: The thermonuclear-burn collaboration compressed work that might have required prolonged coordination across multiple human experts into one focused session, but GPT-5’s premature confidence required expert oversight.The agent also smoothed over thorny issues, creating familiar failure modes that the human had to address.
5 Human-AI synergy in scientific discovery: open questions
The section frames human-agent scientific discovery as an open research problem: understanding when collaboration produces complementary performance and how team utility can be optimized. It proposes a toy utility model and questions spanning environments, trust, over-reliance, and interaction design.
- Motivation: The central question is whether human-agent teams can make discoveries that neither humans nor agents would make alone.This motivates research building on identified risks and case studies.
- Toy model: Team utility U(H, A) is modeled as collaboration advantage CA(H, A) plus collaboration disadvantage CD(H, A).The proposed goal is to maximize CA(H, A) while minimizing CD(H, A).
- Environment factors: Scientific environments may increase utility when outputs are subjective, tasks are long-horizon and open-ended, or agents infer structure from partial observations.These conditions make expert judgment and exploratory sub-task evaluation valuable.
- Environment factors: Automatically verifiable agent trajectories or abundant training data may instead produce environments where CD overwhelms CA, affecting benchmark and application design.Protein structure prediction is given as an example of an environment with abundant training data.
- Scientific teamwork: Open teamwork questions ask how trust, over-reliance, asynchronous communication, and mixed-initiative interaction affect U(H, A), CA, and CD.The section specifically asks how to maintain trust, reduce over-reliance risks, and increase CA without an unreasonable increase in CD.
6 Alternative Views
A central opposing view is that the practical benefits of human-agent systems may be outweighed by the costs and complexity of evaluating them. The authors acknowledge this challenge but propose staged evaluation beginning with lower-cost methods before expensive centaur evaluations.
- The main objection is that evaluation costs and complexity may offset any practical utility gains from human-agent systems in scientific teams.These costs are not captured by collaboration disadvantage CD.
- Staged evaluation could begin with surveys, questionnaires, focused unit-tests, agent-log observations, aggregate statistics, and case studies before investing in centaur evaluations.This approach addresses the open-ended and long-horizon nature of scientific discovery tasks.
7 Conclusion
AI agents in scientific teams should be studied as members of human–agent systems rather than solely as isolated optimizers. This lens emphasizes sustained human participation, intervention throughout discovery workflows, risk mitigation, and human–AI synergy.
- Conclusion: AI agents in scientific teams should be studied as members of human–agent systems, not only as solving isolated optimization problems.The conclusion shifts the unit of analysis from isolated agent performance to participation in a human–agent system.
- Conclusion: The human–agent lens centers sustaining human participation throughout scientific discovery workflows.It frames participation as an ongoing feature of discovery rather than limiting human involvement to artifact review or fixed gates.
- Conclusion: It enables fine- and coarse-grained intervention beyond artifact review and fixed gates, while offering pathways to mitigate near-term risks and cultivate human–AI synergy.The conclusion presents intervention across multiple workflow scales as a means of supporting safer and more synergistic scientific discovery.