Source-linked AI summary
Large Language Models to Accelerate Organic Chemistry Synthesis
Yu Zhang, Yang Han, Shuai Chen, Ruijie Yu, Xin Zhao, Xianbin Liu, Kaipeng Zeng, Mengdi Yu, Jidong Tian, Feng Zhu, Xiaokang Yang, Yaohui Jin, Yanyan Xu
TL;DR
Organic synthesis still relies on costly trial-and-error workflows, motivating AI assistance. This paper develops Chemma, a fine-tuned chemistry LLM integrated with active learning and Bayesian optimization, and reports strong benchmark and experimental results, including 72.2% retrosynthesis top-1 accuracy and a 67% isolated yield for an unreported reaction.
Problem
Chemical synthesis relies on laborious and costly trial-and-error workflows, motivating advanced AI assistants for organic chemistry.
Method
Chemma is a fully fine-tuned LLM integrated with active learning and Bayesian optimization for reaction prediction, exploration, and optimization.
Results
Chemma achieves 72.2% top-1 accuracy in template-free single-step retrosynthesis and supports human–AI exploration of an unreported reaction.
Takeaways & Limitations
Chemma can assist chemical synthesis by extracting reaction knowledge from data and supporting iterative exploration with chemists.
Takeaways & Limitations
Reactions with very limited observations remain difficult, and unreported reactions require wet experiments and fine-tuning before Chemma adapts.
Abstract
from arXiv · showhide
Chemical synthesis, as a foundational methodology in the creation of transformative molecules, exerts substantial influence across diverse sectors from life sciences to materials and energy. Current chemical synthesis practices emphasize laborious and costly trial-and-error workflows, underscoring the urgent need for advanced AI assistants. Nowadays, large language models (LLMs), typified by GPT-4, have been introduced as an efficient tool to facilitate scientific research. Here, we present Chemma, a fully fine-tuned LLM with 1.28 million pairs of Q&A about reactions, as an assistant to accelerate organic chemistry synthesis. Chemma surpasses the best-known results in multiple chemical tasks, e.g., single-step retrosynthesis and yield prediction, which highlights the potential of general AI for organic chemistry. Via predicting yields across the experimental reaction space, Chemma significantly improves the reaction exploration capability of Bayesian optimization. More importantly, integrated in an active learning framework, Chemma exhibits advanced potential for autonomous experimental exploration and optimization in open reaction spaces. For an unreported Suzuki-Miyaura cross-coupling reaction of cyclic aminoboronates and aryl halides for the synthesis of $α$-Aryl N-heterocycles, the human-AI collaboration successfully explored suitable ligand and solvent (1,4-dioxane) within only 15 runs, achieving an isolated yield of 67%. These results reveal that, without quantum-chemical calculations, Chemma can comprehend and extract chemical insights from reaction data, in a manner akin to human experts. This work opens avenues for accelerating organic chemistry synthesis with adapted large language models.
a streamlined closed-loop workflow that leveraged machine learning algorithms to guide robotic
Chemma is presented as a fine-tuned LLM chemistry assistant for reaction prediction, condition generation, performance prediction, and exploration, with wet-experiment validation in open reaction spaces.
- Chemma is a fully fine-tuned LLaMA-2-7b LLM designed as a generative assistant for organic synthesis.
- Chemma answers chemistry questions and supports forward prediction, retrosynthesis, reaction performance prediction, condition generation, and reaction exploration and optimization.
- Chemma’s abilities were assessed with open benchmark data and wet experiments.
- 15 runs were sufficient to explore suitable conditions for an unreported Suzuki–Miyaura cross-coupling synthesis of α-Aryl N-heterocycles.The reported exploration concerned cyclic aminoboronates and aryl halides.
- The results support Chemma’s capability to explore open reaction space.
Design of Chemma as a chemistry assistant
Chemma combines molecular string representations with natural-language interaction, multi-task fine-tuning, ranking data, and regression layers to assist synthesis and reaction optimization.
- Chemma represents molecules with SMILES and integrates this chemical language with natural language for interaction with chemists.
- The assistant covers forward reaction prediction, retrosynthesis, condition generation, and yield or selectivity prediction through chemistry prompts.
- Chemma supports human–AI interaction and an active-learning experiment assistant for reaction exploration and optimization.
- Wet-experiment feedback is incorporated iteratively, after which Chemma is fine-tuned to adapt to the specific reaction.
- The model is trained with multi-task Q&A datasets and fully fine-tuned from LLaMA-2-7b.
- Chemma-RM uses pair-wise ranking Q&A data to identify preferred reaction conditions, while regression networks predict yields or selectivities from reaction embeddings.
Performance of Chemma on open benchmark data
Across benchmark tasks, Chemma reports strong retrosynthesis, ligand recommendation, yield, selectivity, and Bayesian-optimization results, while performance varies with reaction complexity and data coverage.
- Single-step retrosynthesis: 72.2% top-1 accuracy was achieved for template-free single-step retrosynthesis with unknown reaction class.
- Single-step retrosynthesis: 17.1% higher top-1 accuracy than NAG2G (55.1%) was reported for single-step retrosynthesis.
- Ligand recommendation: 15 of 16 base-solvent combinations had recommended ligands with the best median reaction yields.
- Yield prediction: The reported RMSE and R2 for imidazole C–H arylation were 6.59% and 0.74, respectively, while Buchwald–Hartwig results were 6.56% and 0.79.
- Limitations: Prediction performance was weaker for some reactions, possibly because of imbalanced yields, product diversity, and sparse reaction conditions.
- Yield prediction: Chemma-enhanced RF reached R2 values of 0.53 for Suzuki and 0.72 for Buchwald–Hartwig reactions using 5% real data.
- Bayesian optimization: 98.7% was achieved within the first 10 experiments by Chemma-BO for Buchwald–Hartwig, and 99.8% within 25 experiments.
Exploring and optimizing open reaction spaces with Chemma
Chemma is integrated with active learning to explore open reaction spaces through iterative suggestions, wet experiments, feedback, and model adaptation.
- Open reaction-space exploration is motivated by the limits of predefined condition pools and reliance on expert prior knowledge.
- Chemma is positioned as a chemistry assistant within an active-learning framework for open reaction exploration and optimization.
- The workflow iteratively follows Chemma’s suggestions with wet experiments and can fine-tune the model when target yields are not reached.
- For a Pd-catalysed imidazole C–H arylation reaction, Chemma’s suggested ligand reached 76.63% initially, 91.19% with CgMe-PPh, and 100% in the fourth run.
- For a Buchwald–Hartwig reaction, the yield increased from 4.95% with XPhos to 21.66% with tBuXPhos and 84.42% in the 12th run.
- Chemma still faces challenges generating highly reactive ligands for unseen reactions during training, motivating fine-tuning within the active-learning loop.
- In the unreported α-Aryl N-heterocycle reaction, changing the solvent to 1,4-dioxane and selecting PAd3 produced a desired yield of 67%.
- Suitable ligands and solvents were explored within 15 runs, with isolated yields of 45%–67% across additional aryl electrophiles.
Discussion
Chemma is presented as a chemistry-focused assistant that learns from reaction literature, supports multiple synthesis tasks, and can explore open reaction spaces through human–AI collaboration. The discussion also emphasizes limited-data behavior, hallucination, misuse, and the need for expert oversight and further investigation.
- Chemma learns from reactions in the literature and supports retrosynthesis, performance prediction, condition generation, and reaction exploration and optimization.
- Chemma achieves state-of-the-art performance across multiple chemical tasks and can instruct chemical synthesis experiments.
- Reactions with very limited observations remain difficult, and Chemma may generate suboptimal outcomes in such settings.
- For unreported reactions, Chemma initially requires chemists’ feedback, after which iterative wet experiments and fine-tuning support further exploration.
- Chemma may hallucinate non-viable routes or incompatible conditions, so generated protocols should be reviewed by experienced chemists before wet experiments.
- 15 runs yielded 67% for a previously unreported N-heterocycles Suzuki reaction during Chemma-guided exploration of optimized conditions.
Methods
Chemma combines instruction-based generation with specialized training components for reaction tasks and regression-based performance prediction. Its training uses reaction Q&A, preference comparisons, and reaction embeddings to support generation and prediction.
- Chemma uses task-specific instruction prompts to generate potential solutions for chemistry queries such as optimized ligand selection.
- Chemma-SFT fully fine-tunes LLaMA-2, while Chemma-RM predicts effective conditions using preference data tied to reaction performance.
- Preference annotations compare candidate conditions according to reaction performance, such as labeling XPhos as better than tBuXPhos.
- The generation models use cross-entropy objectives, whereas reaction performance prediction uses mean squared error as a regression loss.
- For regression tasks, Chemma extracts reaction embeddings and feeds them into a five-MLP feedforward network.
Implementation of Chemma-BO
Chemma-BO combines Chemma-generated yield predictions with experimental observations and a probabilistic surrogate model to select successive reaction conditions. The loop updates predictions and experiments iteratively until improvement is unlikely or resources are exhausted.
- Bayesian optimization begins with initial reaction-space conditions and uses observed yields to train a probabilistic surrogate model.
- New experiments are sequentially chosen by optimizing an acquisition function for expected utility, then used to update the surrogate posterior.
- The optimization repeats top-five experimental selections until acceptable yields are reached, resources are depleted, or further improvement is improbable.
- Chemma generates yields across varying conditions, after which the top five predicted conditions are selected for experimental validation.
- Chemma-generated yields are combined with Gaussian-process outputs to update predicted yields across the reaction space.
Active learning framework of Chemma for reaction optimization
The active-learning framework uses Chemma to propose conditions in closed or open reaction spaces, tests them experimentally, and incorporates feedback through repeated interaction or model fine-tuning. Optimization ends when acceptable performance is obtained.
- Chemma supports reaction exploration and optimization through an active-learning framework integrated with human–AI interaction.
- The framework handles both closed reaction spaces with predefined conditions and open spaces without experts’ condition limits.
- In round 0, Chemma generates initial conditions using zero-shot prompts, which are then tested in wet experiments to obtain observed yields.
- The process continues through successive exploration and optimization rounds until chemists obtain acceptable performance.
- When experiments provide insufficient performance or data, chemists generate further conditions; with sufficient data but poor performance, they fine-tune Chemma for a new round.
Data and model availability
Chemma's training and testing use open benchmark and experimental datasets, and the retrosynthesis source data are publicly available. The model is also available for free usage.
- Training and testing use the ORD and USPTO open benchmark datasets.
- Retrosynthesis source data are available through Zenodo.
- Yield-prediction data include HTE, ELN, and literature datasets.
- Regioselectivity and enantioselectivity prediction use separate source datasets.
- Chemma is available for free usage.
Municipal Science and Technology Major Project (2021SHZDZX0102), and the Fundamental
The supplied passage identifies research funding from the Central Universities.
- Research funding includes support from the Central Universities.
Inclusion & ethics
The paper states that qualifying contributors are listed as co-authors and other contributors are acknowledged separately.
- Contributors meeting authorship criteria are listed as co-authors.
- Contributors who do not meet all authorship criteria are listed in the Acknowledgements.
Competing interests
The supplied passages describe Chemma's chemistry-assistant functions, training strategy, evaluation tasks, reaction-space exploration, and associated experimental workflows. They also state that the authors declare no competing interests.
- Competing interests: The authors declare that they have no competing interests.
- Reaction-space exploration: The active-learning workflow combines Chemma suggestions, wet experiments, electronic records, and iterative model updating.
- Functions and applications: Chemma supports forward prediction, retrosynthesis, condition generation, and reaction-performance prediction.
- Model and training: The system uses task-specific and ranking prompts, supervised fine-tuning, and experimental feedback for iterative assistance.
- Yield prediction and optimization: Chemma-generated data are evaluated for yield prediction across Suzuki–Miyaura, Buchwald–Hartwig, and imidazole arylation reactions.
- Reaction-space exploration: The framework addresses both expert-defined closed reaction spaces and open spaces not limited by prior expert knowledge.