Source-linked AI summary
Science Done on a Machine by a Machine: AI Agents in Computational Chemistry
Pavlo O. Dral, Hassan Nawaz, Arif Ullah
TL;DR
Computational chemistry simulations are difficult to perform and adoption is constrained by the limited number of experts with the required skills. This Perspective meta-analyzes 49 agentic systems and finds that reported success rates decline as tasks become more demanding, while current systems remain short of autonomous-scientist capability.
Problem
Computational chemistry simulations require specialized expertise, creating a substantial barrier to broader adoption.
Method
The Perspective conducts a meta-analysis of all 49 reported agentic computational-chemistry systems.
Results
99.6% success for a low-autonomy single task declined to 70.9% and 45.6% when tasks demanded more.
Takeaways & Limitations
The surveyed systems collectively include components of machine-performed computational chemistry, but they have not yet reached autonomous-scientist capability.
Takeaways & Limitations
Current agentic systems are not yet autonomous scientists.
Abstract
from arXiv · showhide
We are witnessing an explosion of agentic systems for computational chemistry simulations: from half a dozen in 2024 to a dozen in 2025, and the current number approaches fifty, surveyed in this Perspective as of 8 August 2026. The capabilities of these agentic systems are shifting from assisting in performing a selection of computational tasks to autonomous design and execution of \textit{in silico} experiments, their analysis, and even manuscript writing. The ultimate destination is a fully autonomous AI scientist, where the entirety of computational chemistry is performed on a machine by a machine, without human supervision. While we are not there yet, and all reported systems currently involve a human in the loop, the trend is unmistakable. Even building specialized agentic systems for computational chemistry is increasingly commoditized by generalist agents, which may in the end replace the need for the specialized ones altogether, since adding a new capability will be as easy as asking AI to do it for you. Both the explosion in their number and the very limited adoption beyond their own developers point that way, and we close this Perspective on what it leaves us to do. The speed and scale of disruption agentic systems are bringing to computational chemistry leave many of us dumbfounded about the field's future and what we should spend our efforts on, as already established specialists, teachers, and students, and we have no answer.
1 What is delegated to a machine
Agentic systems aim to shift computational chemistry from automating repetitive calculations to delegating planning, execution, analysis, and writing to machines, reducing dependence on years of specialist training. Their number and capabilities have expanded rapidly, although coverage is broad for structures and dynamics but thin for reaction mechanisms.
- Human expertise: Computational chemistry simulations require highly non-trivial expertise, creating a substantial barrier to broader adoption of best practices.Experts need years of skill development to perform these simulations.
- Human expertise: Automation shifted human work toward ideation, planning, setup, monitoring, analysis, failure handling, and recalculation rather than reducing expert workload.Scripts and workflows automated repetitive tasks, but decision making moved to a higher level.
- Agentic delegation: AI agentic systems promise to make on-the-fly decisions and democratize computational chemistry for researchers without years of specialist training.The envisioned systems perform simulations, analyze results, and write papers.
- Expansion and coverage: Early systems specialized in one or a few tasks, whereas newer systems trend toward broader capabilities spanning structures, molecular dynamics, electronic structure, and periodic density functional theory.Structures are the most common task, while screening, interatomic-potential development, excited states, and free energies have fewer dedicated systems.
- Expansion and coverage: Reaction mechanisms remain thinly represented: transition states and reaction barriers appear in only a few systems, while free energies appear 12 times.The survey characterizes structures and dynamics as broadly covered but reaction mechanism as thinly covered.
2 How much is delegated to a machine
Agentic systems are evaluated by the size of the work they can autonomously complete, progressing from single calculations to experiments, papers, and potentially broader research campaigns. Systems have shifted rapidly toward experiment-scale autonomy, while paper-level capabilities remain limited and campaign- or agenda-level autonomy has not been publicly demonstrated.
- Autonomy scale: Autonomous work is categorized from single calculations, through in silico experiments and paper completion, to campaign-scale investigations and agenda setting.The scale expands from geometry optimization to planning, executing, and analyzing calculation sequences, then to ideation, project management, and multiple papers.
- Autonomy scale: None of the systems publicly reports campaign-level autonomy, although frontier laboratories are already pursuing such efforts.Campaign-level autonomy would involve a series of investigations within a scoped project leading to multiple papers.
- Autonomy scale: Agenda-level autonomy remains premature to discuss before campaign-level autonomy is established, barring a faster arrival of superintelligence.Agenda-level autonomy would define investigations, shape projects, manage resources, and drive them to completion without human intervention.
- Autonomy scale: Agentic systems are becoming more capable and autonomous, shifting from single-calculation operation to experiment-scale capabilities that now constitute the bulk.This trend follows the broader evolution of agentic systems toward increasing autonomy.
- Autonomy scale: Only a few systems claim paper-level capabilities, without clear evidence that they can produce papers en masse without human supervision.Paper-level autonomy is described as the current frontier, but its ability to operate at scale remains unsubstantiated.
3 The evolution of architectures enabling the increas-
Increasing autonomy depends on more mature architectures and stronger underlying AI models, but rapidly changing general-purpose models are not necessarily aligned with autonomous computational chemistry. Architectures have consequently evolved from predefined tool selection toward code, MCP, skills, and loop-based execution, enabling more complex tasks while retaining earlier layers.
- Architectural drivers and constraints: Higher autonomy depends on architecture maturity and underlying AI capabilities, while rapidly changing models may not align with autonomous computational chemistry research.Experts therefore repurpose general-purpose tools for research goals, creating a development lag.
- Architectural evolution: Early systems used predefined tool functions, leaving LLMs mainly to select tools and parameters, with some systems generating input files from software documentation.These tool functions were written before execution and shipped with the system.
- Architectural evolution: Later systems extended capabilities through agent-written executable code, Model Context Protocol calls to external tools, and retrieved skills libraries.Skills provide written procedures that agents follow in place of function calls.
- Orchestration: The orchestration layer shifted from LangGraph-centered early architectures toward general-purpose agents using loosely collected skills.Newer layers accumulate beside predefined functions rather than replacing them.
- Task complexity: Modern systems can work in loops where results change subsequent execution, enabling increasingly complex tasks beyond the predefined sequences typical of initial systems.This capability follows the architectural shift toward extensible, adaptive execution.
4 Where do we stand now and what’s the end game?
Autonomous computational-chemistry tools are becoming more capable and robust, but they still operate within user-defined boundaries and require human supervision, especially for research-scale work. Their rapid commoditization by general-purpose coding agents, limited uptake, and the possibility of machine-performed computational chemistry make evaluation, adoption, and the field’s future difficult to resolve.
- Current capabilities: Autonomous tools are here to stay, increasingly checking their own work and spawning parallel subagents, while experts shift toward supervision.Agents still perform work within boundaries set by users rather than acting as autonomous scientists or obedient calculators.
- Evaluation and reproducibility: Evaluation is increasingly difficult because human judgement is scarce, AI systems evolve rapidly, and leaked benchmarks can contaminate later evaluations.Published claims are further complicated by token costs, undisclosed code or platforms, unclear reuse rights, and systems that may not run as published.
- General-purpose agents and adoption: General-purpose coding agents lowered the barrier to building specialized chemistry systems but are improving faster and raising the barrier to their adoption beyond developers.The authors report that experts may instead adjust general-purpose systems, while their publicly available systems saw slower-than-expected uptake and even group members preferred general-purpose agents.
- The end game: The proposed end game is computational chemistry performed entirely on a machine by a machine, with every component already present somewhere but not yet combined in one system.The authors characterize current agent builders as potentially competing toward their own displacement, expressing this as their opinion rather than an established outcome.
Authors contributions
A.U. conceived and initiated the Perspective, H.N. expanded the draft and surveyed agentic systems, and P.O.D. shaped its scope, database, and final manuscript. All authors revised and approved the work, which was produced with assistance from AI agents under P.O.D.’s direction and verified by the authors.
- A.U. conceived the Perspective and drafted its first version; H.N. extended and revised it and prepared the initial survey of agentic systems.
- P.O.D. set the argument and scope, built the survey database, wrote the final manuscript, and directed subsequent polishing with AI assistance.The authors verified the AI-assisted intermediate versions, figures, and subsequent polishing.
- All authors discussed, revised, reviewed, and approved the manuscript, while the authors stated that their opinions are personal and do not represent their institutions.
Conflict of interest
The authors disclose that Aitomia and Protomia are their own systems and state that both were evaluated under the same rules and sources as all other entries. P.O.D. has financial and authorship ties to Aitomistic and Aitomia, while the other authors declare no competing financial interest.
- Disclosure: Aitomia and Protomia are two surveyed systems owned by the authors, with Protomia used to prepare the article.Protomia’s use is described in Methods.
- Evaluation: Both systems were scored under the same rules as every other entry, using the same public sources and rules stated before any count.The scoring rules were stated in Methods before counts were given.
- Financial interests: P.O.D. co-founded and holds equity in Aitomistic, develops Protomia, and authored Aitomia; the remaining authors declare no competing financial interest.The disclosure identifies P.O.D.’s equity and authorship relationships.
Data availability
No new data were generated, and the survey records are not deposited because they largely reproduce copyrighted text from surveyed papers; the full survey appears in Table 1.
- No data were generated in this work, and the records underlying the survey were not deposited.
- The records largely comprise verbatim quotations copyrighted by publishers, while the complete survey is provided in Table 1.
Methods
The survey assembled a time-bounded, deliberately scoped corpus of computational-chemistry agents using explicit standards for reported tasks and demonstrated delegation. Generative AI supported literature retrieval, database construction, verification, and drafting, while the results remain limited by paper-based evidence, incomplete recall, and subjective placements.
- Corpus construction: The literature search ran through 8 August 2026, combining three survey reference lists, arXiv and ChemRxiv listings, and colleague referrals rather than one reproducible query.Systems were ordered by first public appearance, whether preprint or journal.
- Counting standards: Scientific tasks counted author-reported implementations, including capabilities without worked examples, while excluding future work and merely cited programs.Delegation used the stricter rule of counting only demonstrations in the paper’s body.
- Scope: The survey selected systems rather than tasks, recorded every task performed by included simulation agents, and excluded adjacent domains to preserve specificity to computational chemistry.Excluded areas included retrosynthesis, non-simulation property prediction, robotic synthesis, materials informatics, and general science systems.
- AI-assisted workflow: Generative AI assembled the survey, built the database, drafted intermediate versions, retrieved literature, verified citations, and resolved dates using Claude Code, Claude Opus 5, and Protomia.The final version was written by P.O.D. and polished by the collective.
- Limitations: The results are lower-bound, paper-based measurements: the corpus grew from 37 to 49 systems, and system placements lacked independent recoding.The survey read papers rather than running systems, so demonstrated performance may undercount actual capability.
Appendix: the surveyed systems
The appendix catalogs 49 agentic systems for atomistic simulation surveyed through 8 August 2026, documenting their public appearance dates, reported scientific tasks, and demonstrated delegation. The table is placed outside the main text because of its length.
- Table placement: The table is placed in the appendix rather than the main text because of its length.
- Survey scope: 49 agentic systems for atomistic simulation are surveyed, ordered by their dates of first public appearance through 8 August 2026.The survey cutoff is 8 August 2026.
- Survey criteria: Scientific tasks include capabilities claimed by the authors and implemented in the systems, excluding future work and programs merely cited.A worked example is not required for an implemented capability to be recorded.
- Survey criteria: The Delegation column records what each paper demonstrates in its body.