Source-linked AI summary
Autonomous discovery in the chemical sciences part I: Progress
Connor W. Coley, Natalie S. Eyke, Klavs F. Jensen
TL;DR
Chemical-science discovery lacks a unified framework for comparing automation across matter, processes, and models, while practical and methodological challenges remain. This review classifies discoveries as search problems, proposes questions for evaluating autonomy, and examines case studies across chemical domains. The cases show computer assistance and automation contributing to discoveries, including links identified through literature mining and an active-search accuracy of 80.1% versus 74.0% for a naive strategy.
Problem
The review addresses how automation contributes to different aspects of chemical-science discovery and how the extent of autonomy should be evaluated.
Method
The review classifies discoveries of physical matter, processes, and models as search problems, proposes evaluation questions, and synthesizes case studies across chemical domains.
Results
The case studies report computer-assisted and automated discoveries, including a link between magnesium and migraines and 80.1% model accuracy versus 74.0% for a naive strategy.
Takeaways & Limitations
The cases illustrate how hardware automation and machine learning are transforming experimentation and modelling in the chemical sciences.
Takeaways & Limitations
Autonomous discovery continues to face practical and methodological challenges, and experimental automation remains limiting outside the narrowest design spaces.
Abstract
from arXiv · showhide
This two-part review examines how automation has contributed to different aspects of discovery in the chemical sciences. In this first part, we describe a classification for discoveries of physical matter (molecules, materials, devices), processes, and models and how they are unified as search problems. We then introduce a set of questions and considerations relevant to assessing the extent of autonomy. Finally, we describe many case studies of discoveries accelerated by or resulting from computer assistance and automation from the domains of synthetic chemistry, drug discovery, inorganic chemistry, and materials science. These illustrate how rapid advancements in hardware automation and machine learning continue to transform the nature of experimentation and modelling. Part two reflects on these case studies and identifies a set of open challenges for the field.
2 Introduction
This review frames chemical-science discovery as searches across physical matter, processes, and models, and assesses automation through explicit questions about autonomy. It surveys computer-assisted and automated case studies while recognizing continuing practical and methodological challenges.
- Despite increasingly realistic prospects for automated discovery, universal synthesis platforms remain highly constrained in practice and autonomous discovery still faces primary obstacles.
- Automation and computation improve scientific productivity through efficiency, error reduction, and the ability to address large-scale problems.
- The review classifies chemical-science discoveries into physical matter, processes, and models, unifying them as searches in high-dimensional design spaces.
- It proposes questions for evaluating how much a discovery can be attributed to automation or autonomy, including whether the process is closed loop.
- Case studies span synthetic chemistry, drug discovery, inorganic chemistry, and materials science, covering computer assistance and hardware automation.
3 Defining discovery
The review defines discovery in the chemical sciences across physical matter, processes, and models, and argues that all can be treated as searches through design or hypothesis spaces. It emphasizes that search spaces are usually constrained by goals, domain knowledge, and practical validation requirements.
- Classifications of discoveries: Discovery criteria such as novelty, usefulness, and nonobviousness are inherently subjective and difficult to define precisely.The review notes that “novel” differs from merely “new” and may indicate nonobviousness or lack of predictability.
- Classifications of discoveries: The review distinguishes three discovery types: physical matter, processes, and models.Physical matter includes molecules, materials, and devices; processes include reaction routes and operating conditions; models include empirical, symbolic, natural-law, and conceptual models.
- Models: Models often serve as surrogates for experiments and are produced through fitted empirical relationships, symbolic regression, or more abstract hypothesis searches.Once representations and model families are selected, empirical model fitting can become a defined parameter search, whereas mechanistic explanations occupy harder-to-formalize hypothesis spaces.
- Discovery as a search: Scientific discovery can be framed as a search problem, including searches through chemical, process, and hypothesis spaces.Molecular discovery searches chemical space, process discovery varies process variables or operation sequences, and model discovery searches parameter, mathematical, or hypothesis spaces.
- Discovery as a search: The relevant search space is typically much larger than the region a computer-assisted discovery system is allowed to explore.Restrictions may target scaffolds, material classes, process variables, or hypothesis structures, thereby incorporating domain expertise and simplifying automated validation and feedback.
- Role of validation and feedback: Discovery efforts commonly proceed iteratively through design, making, testing, and revision, or through hypothesize, validate, and revise cycles.Experiments and computational tests support or refute hypotheses, while new information can improve understanding when existing information is insufficient for confident prediction.
4 Elements of autonomous discovery
The review distinguishes automation from autonomy by examining how broadly goals are defined, how searches and experiments are selected, and how results feed back into discovery. It proposes seven questions for assessing autonomy and identifies data, algorithms, and experimental capabilities as complementary enablers.
- Automation versus autonomy: Automation reduces manual laboratory work, but convenience alone may only modestly accelerate discovery rather than fundamentally change its conduct.The distinction between automation and autonomy concerns cognitive burdens such as experimental design and analysis.
- Questions for evaluating autonomy: The framework evaluates autonomy through goal breadth, search-space constraints, experiment selection, brute-force advantage, validation, interpretation, and broader scientific contribution.These questions span the full hypothesis-driven workflow from mission definition through knowledge generation.
- Search and experiment selection: Narrower search spaces and cheaper experiments reduce the human intervention required to select validation experiments.Cost includes experiment time and risk of failure.
- Search efficiency: Active learning may require only 20% of experiments to find an optimum, whereas brute-force comparison is difficult to quantify in continuous or virtually infinite spaces.A fixed-library high-throughput screen is treated as brute-force search.
- Feedback and interpretation: Iterative workflows organize experimental and simulation results as structured information, update prior knowledge, and revise beliefs before the next design round.Compatible output formats make closing the loop a practical step; some specialized workflows naturally select subsequent experiments from results.
- Enabling factors: Data, algorithms, and experiments contribute differently: their integration is needed for fully autonomous discovery.Data support hypothesis generation and refinement, while data combined with algorithms enables virtual screening.
5 Examples of (partially) autonomous discovery
The review surveys case studies across computational reasoning, mechanistic and process discovery, property models, physical matter, and related domains. These examples vary in how automation, data-driven learning, and computational search contribute to discovery.
- Scope of the case studies: The case studies span early computational reasoning, mechanistic models, chemical processes, property models, physical matter, and tangentially related domains.Sections 5.1–5.8 distinguish noniterative and iterative discovery across these categories.
- Different modes of assistance: Some examples primarily automate laboratory hardware, while others learn trends from complex data or explore high-dimensional design spaces computationally.The review emphasizes that the extent of automation and computational assistance differs across cases.
5.1 Foundational computational reasoning frameworks
Foundational programs formalized scientific reasoning through expert rules, heuristics, symbolic regression, and simulated experiments. They reproduced notable laws and models, but fixed priors and limited treatment of uncertainty constrained their utility and discovery scope.
- Compositional models: Early compositional-model programs inferred proposed elements and compounds from lists of chemical reactions using rule-based stoichiometric-like transformations.One example reconstructed aspects of Lavoisier’s theory of oxygen.
- Scope boundary: These programs mainly induced models and generated hypotheses rather than selecting experiments and automating validation or feedback.This marks a boundary between computational reasoning frameworks and closed-loop discovery.
- Rule-based induction: BACON formalized inductive reasoning by searching arithmetic combinations of input variables for regularities in data.BACON.4 applied this framework specifically to chemical problems.
- Symbolic regression: Symbolic-regression systems rediscovered Hamiltonians, Lagrangians, and geometric conservation laws from empirical motion-tracking data.The process generated and scored hypothesized analytical laws.
- Limitations: Such logic frameworks had limited utility because they omitted stoichiometry, phase changes, uncertainty, competing hypotheses, and information requests.The limitation applied despite success formalizing a specific type of scientific reasoning.
- Heuristic discovery: KEKADA used heuristic operators and simulated metabolic experiments to rediscover the Krebs cycle from empirically obtainable data.Its operators included hypothesis generation, problem selection, expectation setting, and confidence modification.
- Role of prior knowledge: Expert-defined rules can be difficult to specify, and overly stringent priors may prevent models from departing far enough from existing theory for substantial discovery.A lack of prior knowledge about allowed reactions may instead permit unlikely hypotheses to be pursued.
5.2 Discovery of mechanistic models
Mechanistic-model discovery combines reaction enumeration, quantum-chemical calculations, expert templates, and molecular dynamics to search large spaces of pathways and reaction channels. These methods identify kinetic mechanisms, transition states, and potentially novel reactions while confronting severe combinatorial growth.
- Search-space challenge: Reaction-mechanism discovery is difficult because even small atomic species generate millions of possible elementary reactions.Heuristics and calculations are needed to direct the search through this combinatorial pathway space.
- Catalytic mechanisms: MECHEM and related approaches enumerate mechanistic steps to rationalize observed catalytic reactions, with ReaxFF guiding searches toward kinetically likely pathways.These methods address multistep catalytic mechanisms.
- Kinetic mechanism generation: RMG uses expert-defined templates to enumerate elementary reactions and estimates rate constants from first-principles calculations and group-additivity rules.The resulting kinetic and thermodynamic parameters support identifying new reactions and pathways.
- Applications: RMG enables exploration of untested fuel additives’ effects on ignition delay, while related work developed detailed pyrolysis kinetic models.These applications use estimated kinetic and thermodynamic parameters to extend mechanism searches.
- Reaction-channel searches: Double-ended searches enumerate potential products and use iterative electronic-structure calculations to identify plausible pathways and estimate transition-state barriers.Single-ended methods instead perturb reactant geometries along reactive coordinates.
- Molecular-dynamics discovery: An ab initio nanoreactor uses molecular dynamics to observe rare events without heuristic reaction coordinates or product-enumeration rules.Periodic compression imparts kinetic energy and promotes collisions over tractable simulation timescales.
- Prospective discovery: These molecular-dynamics approaches could prospectively predict novel reaction types and support development of new synthetic methodologies.The passage presents this as a prospective application rather than a demonstrated general outcome.
5.3 Noniterative discovery of chemical processes
Noniterative chemical-process discovery treats synthesis planning and reactivity prediction as searches through large spaces of pathways, disconnections, and conditions. Computational heuristics, learned models, and high-throughput screening help prioritize experimentally testable candidates, while validation remains essential.
- Synthetic pathway discovery: Synthetic pathways are prerequisites for producing target molecules, and their discovery requires searching combinations of known reactions or generating unseen reactions.The search becomes harder for novel molecules because naive retrosynthetic expansion grows as b^d.
- Synthetic pathway discovery: Expert filters and reaction-rule databases support enumeration of compatible multistep sequences, but pathway validation still requires experiments or expert review.Rule sets range from dozens of expert transformations to thousands or hundreds of thousands of extracted templates.
- Synthetic pathway discovery: CASP programs guide retrosynthetic search with value functions that estimate synthetic complexity and action policies that prioritize transformations.Policies may use nearest-neighbor retrieval or neural-network classification integrated with Monte Carlo tree search.
- Synthetic pathway discovery: Learned value functions can find shorter pathways or optimize user-defined costs, while neural-network-guided search produced pathways chemists judged as plausible as literature routes.The latter result was reported in a double-blind study.
- Reactivity and process discovery: Interpretable models can expose chemical features associated with reactivity or crystallization outcomes, while computational reaction prediction remains useful only with experimental validation.The review also identifies computational chemistry as important for predictive chemical-reactivity models.
5.4 Iterative discovery of chemical processes
Iterative discovery closes the loop between automated experimentation, analysis, and search over process conditions. Flow platforms and other autonomous systems optimize objectives across chemical and materials applications, but practical constraints and multistep interactions limit full closure.
- Discovery of optimal synthesis conditions: Closed-loop experimentation repeatedly selects process conditions, performs reactions, measures outcomes, and uses optimization algorithms to guide subsequent trials.Flow platforms simplify variation of continuous variables and inline sampling of crude product streams.
- Discovery of optimal synthesis conditions: Optimization methods include genetic algorithms, simplex, SNOBFIT, adaptive response surfaces, Bayesian optimization, and reinforcement learning.These methods address continuous variables and can be combined with hardware for discrete choices such as catalysts, ligands, and solvents.
- Discovery of optimal synthesis conditions: Modern machine-learning approaches to reaction optimization have not demonstrated clear advantages over previously used statistical methods.Embedding prior knowledge about regression-surface geometry improved hill-climbing efficiency in one reinforcement-learning routine.
- Discovery of optimal synthesis conditions: Multiobjective process optimization can target knowledge of the Pareto front rather than collapsing competing performance measures into one scalar objective.The Pareto front contains settings where improving one metric requires decreasing another.
- Limitations: Multi-step reactions remain difficult because parameter changes propagate downstream, so workflows often decompose steps or use approximate screening instead of true closed-loop feedback.Kinetic-parameter effects may also be difficult to deconvolute even when rate laws are known.
- Materials applications: Materials-focused closed-loop systems have optimized emission intensity, particle size, crystallization, condensate production, and MOF surface area.Prior data from other MOFs enabled diverse initial experiments before iterative empirical optimization.
- Materials applications: ARES performed up to 100 carbon-nanotube growth experiments per day and used Raman monitoring, a random forest prior, and a genetic algorithm to reach a target growth rate.The prior comprised 84 expert-defined experiments.
5.5 Noniterative discovery of structure-property models
Noniterative discovery of structure–property models uses statistical, machine-learning, symbolic, and mechanistic approaches to relate chemical structure or composition to properties, reactivity, spectra, and electronic energies. Interpretability can inform design, while explanations and surrogate models retain important limitations.
- Structure–property models: Structure–property models learn relationships between molecular or material features and properties from data, and interpretable models can inform design.QSAR/QSPR models can represent a belief about a performance landscape for a discovery task.
- Important molecular features: PROGOL identified criteria for mutagenicity from hypothesized toxicophores defined by connectivity and partial charge values.Related applications pursued explanations for carcinogenicity and ACE-inhibition activity.
- Important molecular features: Interpretability methods include few-parameter regressions, decision trees, feature selection, symbolic-rule extraction, and identification of relevant training examples.Random-forest ensembles may obscure descriptor-level analysis.
- Important molecular features: Fragment-contribution explanations estimate feature effects by masking molecular substructures and observing changes in predicted properties.Per-atom or per-substructure importance can oversimplify nonlinear learned relationships.
- Spectral analysis: DENDRAL combined mass-loss and fragment matching with structure enumeration and spectral prediction to propose molecular structures from MS data.Its rapid calculations still required expert heuristics, including a badlist, to prune unrealistic structures.
- Potential energy surfaces and functionals: Surrogate models replace expensive electronic-structure calculations or selected components, including neural-network energy potentials trained on DFT and refined with CCSD(T)/CBS data.Active learning can strategically acquire costly training data.
5.6 Noniterative discovery of new physical matter
Noniterative discovery uses predefined search spaces, virtual screening, or high-throughput experimentation to identify promising physical matter without feeding validation results back into the model. These approaches can generate diverse candidates and practical discoveries, but constrain serendipity and often leave validation or purification as bottlenecks.
- Strategy: Noniterative workflows screen predefined spaces exhaustively or use surrogate models to approximate structure–function landscapes.QSAR/QSPR models trained on experimental or simulated data can rank large candidate sets, but validation results do not revise the model.
- Strategy: Fixed computational maps leave little room for serendipity because compounds predicted as unuseful are generally not tested.Random exploration must be explicitly included to counter this selection bias.
- Drug discovery: DNA-encoded libraries reach hundreds of compounds per well and theoretical complexities of hundreds of millions or billions of compounds.These scales exceed traditional single-compound-per-well high-throughput approaches by several orders of magnitude, although analysis and purification can limit throughput.
- Drug discovery: DNA-encoded and diversity-oriented synthesis strategies have supported discovery of RIP1 kinase and sEH inhibitors, lead compounds, and biological probes.Diversity-oriented synthesis generates structurally and functionally diverse small-molecule collections through varied reagents, starting materials, and multicomponent reactions.
- Materials science: High-throughput materials platforms vary composition, stoichiometry, and deposition sequence to search broad spaces across catalysts, MOFs, coatings, polymers, and alloys.One reported campaign produced the first four lead-free thin-film compositions.
- Molecular generation: Computational molecular-generation methods produce libraries enriched for high-affinity ligands and optimize molecules against objectives such as logP and synthetic accessibility.Graph-based generation addresses representational problems associated with multiple valid SMILES strings.
5.7 Iterative discovery of new physical matter
Iterative discovery uses experiment results to select subsequent experiments and update structure–function models, concentrating measurements on informative or promising regions. Active learning, Bayesian optimization, evolutionary algorithms, and automated platforms expand this feedback-driven search, while practical constraints still limit autonomy.
- 5.7 Iterative discovery of new physical matter: Iterative workflows update structure–function models with experimental results to improve subsequent predictions and select informative or promising experiments.Active learning and Bayesian optimization balance information gain with candidate performance within the search space.
- 5.7.1 Discovery for pharmaceutical applications: Eve screens more than 10,000 compounds per day, builds a surrogate structure–activity model, and selects compounds by predicted activity, selectivity, or prediction variance.Its success depends on a large library containing at least one acceptably high-performing compound.
- 5.7.1 Discovery for pharmaceutical applications: Active learning is typically more economical than brute-force screening because it accounts for hit utility, reduced experimental burden, and missed-hit costs.These factors are incorporated in an econometric model of the drug-discovery process.
- 5.7.1 Discovery for pharmaceutical applications: On-demand synthesis expands chemical search spaces beyond in-stock compounds, including a microfluidic platform producing 27×10 compounds through one-step Sonogashira coupling.The platform integrated synthesis with purification, dilution, and kinase activity assays, while a random forest guided experiment selection.
- 5.7.1 Discovery for pharmaceutical applications: Closed-loop synthesis, purification, and testing remains a proof of concept when the explored design space is extremely narrow and brute-force search may be faster.More flexible integrated platforms have generally been applied to open-loop discovery with manual compound design and synthesis planning.
- 5.7.1 Discovery for pharmaceutical applications: Evolutionary and genetic-algorithm strategies expand beyond pre-enumerated candidates by mutating high-performing compounds under constrained transformations.Related approaches have identified sub-micromolar inhibitors and remained effective for organometallic-complex discovery.
- 5.7.2 Discovery for materials applications: Automated materials searches discovered NiTi-based shape-memory alloys with low thermal hysteresis and BaTiO3-based piezoelectric materials with large electrostrains.Thin-film casting broadened the problems accessible to autonomous platforms.
5.8 Brief summary of discovery in other domains
Computer assistance extends discovery beyond physical matter to genomics, protein engineering, gene–enzyme relationships, and correlations extracted from scientific text. These applications generate hypotheses, guide experiments, and organize biological knowledge, while some systems retain manual operational steps or embed the knowledge they recover.
- Scope: Automated discovery spans many additional machine-learning and experimental applications, but the review presents them as a broad set of approaches rather than a complete account.The authors point readers to a more comprehensive collaborative review for additional examples.
- Text mining: ARROWSMITH identified overlapping terms in MEDLINE abstracts and helped generate testable hypotheses, including a magnesium–migraine link.Literature mining also serves to organize biological data for automated analysis or manual inspection.
- Genomics: Probabilistic graph models combine genetic, protein-interaction, and metabolic-pathway information to propose hypotheses about gene functions.The approach exploits the large quantity of structured genetic information produced by sequencing advances.
- Protein engineering: Supervised machine learning can assist directed evolution by selecting protein mutants across multiple rounds in high-dimensional, discontinuous structure–function landscapes.These methods are particularly relevant when mutational effects are nonadditive.
- Gene–enzyme relationships: Adam hypothesized enzyme assignments for 15 yeast open reading frames and selected growth experiments involving knockout mutants and metabolites.The demonstration still required manual plate transfer between the liquid handler, incubator, and plate reader.
- Evaluation: Adam’s reported active-search accuracy was 80.1%, compared with 74.0% for a strategy choosing the cheapest unperformed experiment.The authors also acknowledge criticism that the new knowledge was implicit in the problem formulation.
6 Conclusion
The review classifies discovery of physical matter, processes, and models as search problems and proposes questions for evaluating autonomy. Case studies show substantial progress, but true closed-loop discovery remains uncommon outside narrow design spaces and practical and methodological challenges persist.
- Conclusion: The review defines discovery categories as physical matter, processes, and models, unified through a search-problem perspective.It proposes evaluating autonomy through questions about goals, search-space constraints, experiment selection and execution, interpretation, and broader knowledge contribution.
- Conclusion: The proposed autonomy assessment asks how broadly goals are defined, how constrained the search is, and how experiments and results are handled.It also asks whether navigation improves on brute force and whether outcomes contribute to broader scientific knowledge.
- Conclusion: Case studies demonstrate substantial progress toward autonomous discovery, yet few examples achieve true closed-loop discovery.The scarcity is especially evident beyond the narrowest design spaces.
- Conclusion: Practical and methodological challenges remain in the pursuit of autonomous discovery.The review’s second part examines selected case studies and describes remaining challenges requiring further development.