Source-linked AI summary
Dynamic language model representations for multi-objective reaction optimisation
Joshua W. Sin, David Ming Segura, Bojana Ranković, Siu Lun Chau, Marius D. R. Lutz, Andrea Anelli, Ryan P. Burwood, Kurt Püntener, Maximilian J. Notheis, Raphael Bigler, Philippe Schwaller
TL;DR
Reaction optimisation needs representations that remain useful across chemically diverse components and multiple objectives, but descriptors and one-hot encodings have important limitations. The paper learns task-adaptive representations from textual reaction conditions using language models jointly trained with Gaussian-process surrogates, achieving faster convergence than established baselines and successful prospective scale-up in heterogeneous reaction spaces.
Problem
Existing reaction representations are either chemically uninformative or difficult to transfer across chemically distinct components, making a shared representation for multi-objective optimisation challenging.
Method
The LLM-GP framework encodes textual reaction conditions with a LoRA-fine-tuned language model and jointly trains objective-specific Gaussian-process surrogates within Bayesian optimisation.
Results
Across sequential and parallel cross-coupling benchmarks, LLM-GP reached convergence in fewer experiments than descriptor libraries or one-hot encoding; prospective campaigns achieved 94% and 84% isolated yield, with the latter at 99.6% enantiomeric excess.
Takeaways & Limitations
Dynamic language-model representations provide a general route to multi-objective optimisation across chemically diverse reaction systems without computing descriptors for each new system.
Takeaways & Limitations
Descriptor methods remain chemically interpretable, whereas the paper’s approach avoids descriptor selection and computation rather than providing that interpretability.
Abstract
from arXiv · showhide
Optimising chemical reactions across multiple objectives, such as yield, selectivity, and safety, is central to chemical synthesis, and model-driven approaches depend critically on how reaction components are represented. Established featurisations are either chemically uninformative, as with one-hot encodings, or, as with molecular descriptors, do not readily extend across chemically distinct components. For structurally and functionally diverse components, it is therefore unclear what a shared representation should contain. Constructing such a representation is itself a challenging research undertaking that must be revisited for each new reaction system. Here we bypass this step by learning the reaction representation dynamically from text. Textual descriptions of reaction conditions are encoded by a fine-tuned language model trained jointly with Gaussian process surrogates, yielding task-adaptive representations within a multi-objective Bayesian optimisation loop. Across nickel- and palladium-catalysed cross-couplings in both sequential and parallel experimentation regimes, this approach reaches optimisation convergence in fewer experiments than descriptor libraries or one-hot encoding. Applied prospectively to a palladium-catalysed cyanation spanning mixed ligand denticity and heterogeneous additives, and to a three-objective asymmetric hydrogenation across chiral iridium and ruthenium catalyst families, two rounds of high-throughput experimentation (192 reactions, under 3% of each design space) delivered conditions translating directly to gram scale in 94% and 84% isolated yield, the latter at 99.6% enantiomeric excess.
1. Introduction
Reaction optimisation must search large combinatorial spaces while balancing multiple objectives, but existing reaction representations struggle to provide a transferable shared space across chemically diverse components. The paper introduces a dynamically learned language-model representation and applies it to benchmark and prospective optimisation campaigns.
- Reaction optimisation balances yield, selectivity, cost, and safety while navigating vast combinatorial spaces of ligands, catalysts, bases, and other components.
- Descriptor methods require system-specific chemical intuition and may not transfer across distinct components, making feature engineering increasingly difficult as chemical diversity grows.Descriptors can be computationally expensive and may be only partially applicable across ligand classes or heterogeneous additives.
- One-hot encoding is broadly applicable but produces sparse, high-dimensional representations that encode no chemical similarity.
- LLMs offer a unified alternative because their embeddings can encode complex chemical information from text into fixed-length representations across component classes.
- The LLM-GP framework fine-tunes textual reaction representations jointly with objective-specific Gaussian-process surrogates, adapting embeddings during multi-objective Bayesian optimisation.
- Across sequential and parallel cross-coupling benchmarks, LLM-GP converged in fewer experiments than descriptor libraries and one-hot encoding, while prospective campaigns reached gram-scale yields after two 96-experiment rounds.The campaigns achieved 94% and 84% isolated yield, with the latter reaching 99.6% enantiomeric excess, while sampling under 3% of each design space.
2. Computational results
The computational evaluation compares LLM-GP with descriptor libraries and one-hot encoding using hypervolume convergence across sequential and batched optimisation regimes. LLM-GP consistently reaches practical performance thresholds in fewer experiments, with the strongest reported sequential gains on the nickel-catalysed benchmark.
- Hypervolume measures the objective-space volume dominated by the current Pareto-optimal set, while 90% and 95% thresholds quantify convergence speed.
- Sequential low-data reaction optimisation: The sequential benchmarks used yield and selectivity objectives over large categorical search spaces for nickel- and palladium-catalysed Suzuki couplings.The campaigns contained 384 and 404 experiments sampled from spaces of 88,000 and 59,040 combinations, respectively.
- LLM-GP consistently reached practical convergence thresholds in fewer experiments than DFT descriptor and one-hot baselines across the evaluated optimisation settings.
- Sequential low-data reaction optimisation: Approximately 13 iterations reached 90% hypervolume for nickel-catalysed Suzuki coupling with LLM-GP, versus approximately 26 for DFT descriptors and 24 for one-hot encoding.At 95% hypervolume, LLM-GP required approximately 24 iterations, compared with approximately 28 and 30 for the two baselines.
- Sequential low-data reaction optimisation: Approximately 65 iterations reached both 90% and 95% hypervolume on palladium-catalysed Suzuki coupling with LLM-GP, versus approximately 80 for each baseline.
- Parallel high-throughput optimisation: In batched high-throughput optimisation, reducing plate iterations can save experimental time and resources because each 96-well campaign typically requires approximately one week.
3. Prospective experimental case studies
Prospective campaigns tested the framework on chemically heterogeneous cyanation and asymmetric hydrogenation spaces, where unified reaction representations are difficult to construct. Two-round HTE campaigns identified scalable cyanation and hydrogenation conditions across broad multi-objective design spaces.
- Case study 1: Palladium-catalysed cyanation: The cyanation search combined 40 palladium catalysts, three cyanide sources, five additives, ten solvents, co-solvents, and temperatures after excluding conditions above solvent boiling points.The resulting design space contained 30,000 conditions and included mono- and bidentate ligand classes.
- Case study 1: Palladium-catalysed cyanation: Cyanation optimisation initially favoured toxic cyanide sources, but constraining the second round to K4[Fe(CN)6] produced multiple conditions exceeding 99% conversion and selectivity.High-performing hits spanned both monodentate and bidentate ligands.
- Case study 1: Palladium-catalysed cyanation: 94% isolated yield was achieved at gram scale for palladium-catalysed cyanation after two HTE rounds using a lower-hazard cyanide source.The campaign evaluated 192 experiments across a 30,000-condition design space and identified conditions exceeding 99% conversion and selectivity.
- Case study 2: Asymmetric ketone hydrogenation: The hydrogenation search jointly varied 32 chiral iridium and ruthenium catalysts, nine bases, seven solvents, base loading, and temperature across 8,064 conditions.Conversion, syn diastereomeric excess, and syn enantiomeric excess were maximised simultaneously.
- Case study 2: Asymmetric ketone hydrogenation: Syn-selective hydrogenation outcomes increased from 5/96 in Plate 1 to 59/96 in Plate 2 as optimisation refocused on syn-producing regions.The best stereoselective conditions reached 87.1% de (syn) and 93.2% ee (syn), while the balanced lead achieved >99% conversion, 80.5% de (syn), and 89.7% ee (syn).
- Case study 2: Asymmetric ketone hydrogenation: Two HTE rounds delivered asymmetric hydrogenation conditions with 84% isolated yield and 99.6% ee at gram scale.The campaign used 192 experiments from an 8,064-condition design space.
4. Outlook
The framework combines text-based reaction representations with multi-objective Bayesian optimisation to avoid reaction-specific descriptor engineering. Its shared language-model embeddings are jointly adapted with objective-specific Gaussian-process surrogates.
- 4. Outlook: Dynamic language-model representations offer a general route to multi-objective optimisation across chemically diverse reaction systems without selecting or computing descriptors for each new system.The approach encodes ligands, additives, and catalysts directly from textual descriptions, allowing optimisation before system-specific mechanistic understanding is established.
- 4. Outlook: The benchmark compares one-hot encoding, molecular descriptors, and language-model representations as alternative reaction-condition featurisations.One-hot vectors concatenate binary indicators, while descriptor baselines use established libraries where coverage exists and one-hot encoding for uncovered component classes.
- 4. Outlook: Each reaction condition is converted into a pooled language-model embedding from a textual prompt describing component values.Pooling produces a fixed-dimensional representation from token-level hidden states, with architecture-specific mean, last-token, or CLS pooling evaluated.
- 4. Outlook: The optimisation loop begins with Sobol coverage, repeatedly refits surrogates, and selects batches using qLogNParEGO over the remaining design space.Fixed-representation baselines use optimised Gaussian-process configurations, while the learned-representation model jointly optimises language-model and Gaussian-process parameters; experiments use 20 random seeds.
- 4. Outlook: Objective-specific projection heads map the shared embedding to separate Gaussian-process inputs, while gradients jointly fine-tune the language model for predictive performance across objectives.Each projection head receives gradients only from its corresponding objective, whereas the shared language-model parameters receive gradients from all objectives.
Competing interests
Several listed authors declare potential financial and non-financial conflicts of interest as employees of F. Hoffmann-La Roche Ltd.
- Several authors declare potential financial and non-financial conflicts of interest as full employees of F. Hoffmann-La Roche Ltd.
A. Supplemental main text benchmarking studies
The supplemental benchmarks compare LLM-GP with DFT descriptors and one-hot encoding in sequential and high-throughput optimisation settings.
- The sequential benchmark datasets comprise nickel-catalysed and palladium-catalysed Suzuki couplings.
- Sequential benchmarks compare LLM-GP, DFT descriptor databases, and one-hot encoding on nickel- and palladium-catalysed Suzuki coupling datasets.Performance is assessed using iterations to reach 90% and 95% hypervolume thresholds and the percentage of runs reaching those thresholds within budget.
- HTE benchmarks compare the same representations on palladium-catalysed sulfonamide coupling across 24-, 48-, and 96-well plate batch sizes.The evaluations report plate iterations to threshold and the percentage of optimisation runs reaching each threshold within budget.
B. Influence of different pre-trained models and pooling strategies
Supplemental ablations examine how language-model architecture, pooling strategy, and HTE plate batch size affect optimisation performance across coupling benchmarks.
- Influence of different pre-trained models and pooling strategies: The ablations measure optimisation iterations required to reach 90% and 95% hypervolume thresholds across nickel- and palladium-catalysed Suzuki coupling datasets.
- Influence of different pre-trained models and pooling strategies: T5-base with mean pooling consistently reached practical 90% and 95% hypervolume thresholds in the fewest iterations across both Suzuki coupling datasets.The comparison covers T5-chem, BART-base, T5-base, T5-small, and Qwen2.5-0.5B with the pooling strategies specified in the figure.
- Influence of different pre-trained models and pooling strategies: Run-based evaluations report the percentage of 20 random-seed optimisation runs reaching 90% and 95% thresholds within the allocated budget.
- Influence of different pre-trained models and pooling strategies: HTE ablations evaluate mean plate iterations to threshold across 24-, 48-, and 96-well batch sizes.
C. Design spaces for prospective optimisation campaigns
The prospective campaigns explore broad reaction-condition spaces for palladium-catalysed cyanation and asymmetric ketone hydrogenation.
- The prospective Pd-catalysed cyanation space contains 40 mono- and bidentate Pd catalysts, 10 solvents, 3 cyanide sources, 2 co-solvents, 5 additives, and 3 temperatures.
- The prospective asymmetric ketone hydrogenation space contains 32 chiral catalysts, 7 solvents, 9 bases, 2 base mol%, and 2 temperatures.
- The supplemental material identifies the paper as Dynamic language model representations for multi-objective reaction optimisation.
1 High-throughput experimentation (HTE) platform
The HTE platform combined automated parallel experimentation with controlled dispensing and analysis infrastructure. Its glove-box-enclosed setup supported handling of solid and liquid reaction components.
- Platform architecture: The HTE system comprised two interconnected Big Kahuna platforms, with one integrated into a LiCONiC LiCotel system.The setup was enclosed in an LC Technology Solutions glove box with separate circulation systems for solid dispensing and reaction execution.
- Component handling: Solid precursors, ligands, and additives were dispensed through automated hopper systems, while liquids were manually added by pipette.Materials below 0.4 mg were dispensed as coated ChemBeads.
- Platform architecture: Supplementary Figure 1 provides a schematic overview of the high-throughput experimentation setup.
2 Prospective application: Palladium-catalysed cyanation
The prospective cyanation campaign used a palladium design space spanning diverse catalysts, additives, solvents, and temperatures, followed by gram-scale translation. The product was characterized by 1H and 13C NMR data.
- HTE procedure: The HTE procedure dispensed palladium catalysts, cyanide sources, and solid additives into 96-well plates before adding substrate, solvents, cosolvent, and applicable NEt3.Plates were sealed and stirred at 400 rpm for 20 hours at the specified temperature.
- Gram-scale translation: The gram-scale cyanation used DMC/water, K4Fe(CN)6, and [Pd(allyl)(Xantphos)]Cl at 75 °C for 4 hours.The reaction used 1.0 mol% palladium catalyst and 0.50 equivalents of K4Fe(CN)6·3H2O.
3 Prospective application: Asymmetric ketone hydrogenation
The asymmetric ketone hydrogenation campaign used an automated 96-well workflow with iridium and ruthenium catalysts and variable bases. Reactions were prepared under inert atmosphere and processed in a screening pressure reactor.
- HTE procedure: The HTE plate contained 1 mol% iridium or ruthenium catalysts and solid bases dispensed into 1 mL vials.The reactions used stock solutions of 1,2-diphenylpropan-1-one and optional liquid bases.
- Reaction setup: 1,2-Diphenylpropan-1-one stock solution was added to the vials before applicable liquid bases.The stock solution contained 20 mg of substrate in 200 µL solvent.
- Reaction execution: The sealed 96-well plate was placed in a screening pressure reactor and shaken overnight.
O OH [Ir(Cl)H2((S)-DTB-SpiroPAP-3-Me)] (1 mol%)
The gram-scale asymmetric hydrogenation used a chiral iridium catalyst with tert-amyl alcohol and DBU, followed by chromatographic purification and crystallisation. The study also tested model sensitivity to removing syn-selective initial observations and augmented data using catalyst-enantiomer symmetry.
- Gram-scale translation: The gram-scale reaction used chiral iridium catalyst, tert-amyl alcohol, DBU, and hydrogen under autoclave conditions.The catalyst loading was 0.01 equivalents, with 19.2 equivalents of tert-amyl alcohol and 0.5 equivalents of DBU.
- Gram-scale translation: Gram-scale hydrogenation afforded 1.68789 g of purified product, corresponding to 83.6% yield.The product was purified by flash chromatography and crystallised from pentane.
- Data augmentation: Catalyst-enantiomer symmetry was used to augment training data because enantiomeric catalysts reverse ee while preserving conversion and diastereomeric excess.
- Model sensitivity to initial training data: The model recovered conditions reaching 83.4% syn de and 93.3% syn ee even after all five syn-selective initial conditions were removed.The ablation left no positive diastereoselectivity examples in the training data, and evaluated values were lower bounds on recoverable performance.