Source-linked AI summary
Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries
Henrik Wille, Luis-Finley Schütz, Felix Strieth-Kalthoff
TL;DR
It remains unclear how well pretrained molecular representations support iterative discovery across chemical domains. This study benchmarks molecular language models and fingerprints across virtual libraries and finds that domain adaptation improves language-model representation performance, while fingerprints remain robust.
Problem
It remains unclear how well pretrained molecular representations support iterative molecular discovery within and beyond their pretraining domains.
Method
The study benchmarks molecular language models and fingerprints in iterative virtual-library optimization, then evaluates three domain-adaptation strategies using target-library structures.
Results
Molecular language-model embeddings vary considerably across libraries, while fingerprints perform robustly; adapted ChemBERTa, MolFormer, and T5Chem show positive effects relative to fingerprints.
Takeaways & Limitations
Molecular representation quality depends on the target library, and explicit adaptation can improve the practical utility of molecular language-model embeddings for discovery.
Takeaways & Limitations
The benchmark is limited to libraries with computed target properties available for all molecules and treats seed populations from different libraries as unrelated.
Abstract
from arXiv · showhide
Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug discovery, organic materials, and catalysis. Native molecular language model embeddings show substantial variation in discovery performance across libraries, whereas molecular fingerprints provide a consistently strong and robust baseline. Consistent with a potential domain-representation mismatch, we show that explicit domain adaptation substantially improves representation performance. Fine-tuning molecular language model encoders on structures from the target virtual library consistently improves sample efficiency, with several adapted encoders emerging as the top-performing representations across the benchmark tasks. These results show that molecular representation quality depends strongly on the target domain and that explicit adaptation can improve the practical utility of molecular foundation models. More broadly, our findings establish domain-adapted molecular representations as a promising strategy for sample-efficient adaptive decision making in virtual screening and self-driving laboratories.
Introduction
Rapid discovery increasingly relies on data-driven closed design–make–test–analyze loops, where molecular representations are central to learning structure–property relationships from limited evaluations. Because libraries across drug discovery, materials chemistry, and catalysis occupy distinct regions of chemical and representation space, pretrained representations may inadequately distinguish candidates outside their pretraining domain.
- Motivation: Rapid functional-molecule discovery spans medicinal, agricultural, and materials chemistry and increasingly uses data-driven decision-making to close discovery loops.Automated laboratories can access combinatorial libraries when candidate synthesis uses a small number of well-defined procedures.
- Motivation: Molecular representation is central to learning and generalizing structure–property relationships when virtual and experimental discovery operate under limited evaluation budgets.These settings require models that learn effectively from limited property evaluations.
- Representation–domain mismatch: Pretrained embedding spaces are shaped primarily by their represented chemical-space region, while materials and catalysis libraries may occupy different regions and inadequately differentiate candidates.This representation–domain mismatch can limit learning structure–property relationships, particularly near property cliffs.
- Representation–domain mismatch: Representation quality should be judged by how effectively it supports iterative molecular discovery, because distribution shift or embedding contraction alone does not establish inadequacy.The introduction frames practical discovery performance—not distributional change alone—as the decisive criterion.
- Benchmark scope: Benchmark experiments cover six virtual libraries spanning drug discovery, materials chemistry, and catalysis, providing distinct target domains for evaluating molecular representations.The drug-discovery tasks include large ZINC15-based libraries with docking scores for AmpC β-lactamase and 8-oxoguanine DNA glycosylase targets.
- Benchmark scope: Chemical-space analysis confirms systematic distribution differences among the target libraries relative to a 25,000-molecule ZINC15 reference set representing drug-like pretraining chemical space.The reference set was used to investigate distribution shifts and possible representation–domain mismatches.
Native molecular language model representations show variable performance across tasks.
Native molecular language model representations showed variable discovery performance across tasks, with broadly consistent ordering between virtual screening and experimental discovery but less pronounced differences under lower experimental budgets. Explicit domain adaptation mitigated underperformance relative to Morgan fingerprints.
- Discovery regimes: 20,000 property evaluations defined the virtual-screening regime, whereas experimental discovery used batches of 50 up to 1,000 evaluations.These regimes approximated computational campaigns and self-driving laboratory constraints, respectively.
- Native representations: Representation ordering was largely consistent across virtual screening and experimental discovery, although differences were less pronounced in the experimental regime.Under the lower experimental budget, Morgan fingerprints and MolFormer and T5Chem embeddings showed overlapping global posterior effects.
- Domain adaptation: Adapted ChemBERTa, MolFormer, and T5Chem embeddings showed positive global effects relative to Morgan fingerprints across several libraries and acquisition-budget regimes.ChemBERTa benefited most strongly: adaptation substantially reduced or reversed its native embeddings’ significantly negative average effects relative to Morgan fingerprints.
- Representation–domain mismatch: Pretrained representation geometry can substantially mismatch target-library structure, particularly when materials and catalysis libraries occupy chemical-space regions distinct from drug-dominated pretraining corpora.Embedding distributions were shifted and comparatively narrow, but these differences explained little of the observed benchmark-performance variation.
Supplementary Information · Domain-Adapted Molecular Language Models
The supplementary information identifies the paper’s authors and their University of Wuppertal affiliations. It is associated with the work titled “for Efficient Search of Make-on-Demand Libraries.”
- for Efficient Search of Make-on-Demand Libraries: Henrik Wille is listed as an author.
- for Efficient Search of Make-on-Demand Libraries: Luis-Finley Schütz is listed as an author.
- for Efficient Search of Make-on-Demand Libraries: Felix Strieth-Kalthoff is listed as an author.
- for Efficient Search of Make-on-Demand Libraries: The authors are affiliated with the University of Wuppertal’s School of Mathematics and Natural Sciences.The listed address is Gaußstr. 20, 42119 Wuppertal.
- for Efficient Search of Make-on-Demand Libraries: Felix Strieth-Kalthoff is additionally affiliated with the University of Wuppertal’s Interdisciplinary Center of Machine Learning and Data Analytics.The listed address is Gaußstr. 20.
1 Experimental Setup
The benchmark evaluates batch-wise iterative optimization over virtual molecular libraries with fully available target properties, comparing heuristic representations and four pretrained molecular language models, including domain-adapted variants. Performance is assessed through repeated optimization trajectories and normalized AUOC-based statistical comparisons.
- Iterative optimization: Virtual-library optimization begins with a random seed batch, trains a surrogate on cached representations, and iteratively selects batches using acquisition values until budget B is exhausted.Target-property observation corresponds to lookup in the full library, and representation vectors are precomputed before optimization.
- Experimental design: Each experiment was repeated 20 times with different seed datasets, using identical initial datasets across methods to enable paired comparisons.Predictions and acquisition values are recomputed for all remaining library molecules after each observed batch.
- Evaluation: Quantitative evaluation records recall of top-1000 or top-1% candidates versus observations, while normalized AUOC is computed by trapezoidal integration and analyzed with hierarchical Bayesian regression.The regression separates library, model, and seed effects in the paired design.
- Molecular representations: Four molecular language models pretrained on large SMILES corpora provided embeddings from canonicalized SMILES, with pooling used for ChemBERTa, MolFormer, and T5Chem.SmiTed embeddings were obtained from the autoencoder latent space.
- Domain adaptation: Domain adaptation adds a one-hidden-layer projection head and fine-tunes it with a LoRA module applied to the original language-model weights.Unless otherwise noted, fine-tuning used Adam at 5.0 ⋅ 10–4 for 25 epochs.
2 Molecular Library Distributions
The study analyzed six virtual molecular libraries with ground-truth target-property labels and systematically characterized their molecular distributions. Distribution analysis combined fingerprint and transformer-embedding similarities with projections against a 25,000-molecule ZINC15 reference set.
- Scope: Six virtual libraries with ground-truth target-property labels were systematically analyzed for their underlying molecular distributions.The libraries span docking, organic solar-cell, organic-laser, and enantioselectivity applications.
- Distribution analysis: Distribution analysis used Morgan fingerprints and transformer encoders to quantify within-library similarity dispersion.Similarity was computed from 5M random molecular pairs, using Tanimoto similarity for fingerprints and cosine similarity for transformer embeddings.
- Reported outputs: Figures S1–S24 report similarity histograms, dimensionality-reduced projections, and target-property histograms, while Tables S1–S6 provide distribution characteristics.The reported target properties include docking scores, DFT-computed power conversion efficiencies, DFT-computed emission oscillator strengths, and predicted ΔΔG‡ values.
3 Benchmarking Surrogate Models and Acquisition Policies
The section benchmarks surrogate models and acquisition strategies on MolFormer embeddings across virtual-screening and experimental-discovery settings. Evaluation uses optimization trajectories, relative logit-scale effects, and effect ratios across molecular libraries.
- Surrogate models: Surrogate models are compared over 20,000-experiment virtual-screening campaigns with batch sizes of 1,000, using means from 20 independent campaigns.Optimization trajectories include standard-error shading.
- Surrogate models: Surrogate-model effects are evaluated relative to a Laplace Neural network across molecular libraries using 8,000 posterior samples.Effect ratios are also averaged over all libraries with 95% highest-density intervals.
- Acquisition strategies: Acquisition strategies are compared with a Laplace Neural Network surrogate on MolFormer embeddings over 20,000 experiments with batches of 1,000.Optimization trajectories summarize 20 independent campaigns and show standard-error shading.
- Acquisition strategies: Acquisition-strategy effects are measured relative to Top-K acquisition with a UCB acquisition function across molecular libraries using 8,000 posterior samples.Effect ratios are averaged over all libraries and displayed with 95% highest-density intervals.
4 Evaluating Molecular Representations · 5 Supplementary References
The supplementary evaluations compare molecular encoders, pooling strategies, and fine-tuning approaches across virtual screening and experimental discovery budgets. Results are summarized through optimization trajectories, relative logit-scale effects, and effect-ratio distributions, with repeated campaigns and uncertainty estimates.
- 4.1 Heuristic Representations and Pre-Trained Transformer Embeddings: Domain-specific and molecular language model encoders are evaluated with Laplace Neural Network surrogates over 20,000-experiment optimization campaigns.Trajectories average 20 independent campaigns, use batch sizes of 1,000 experiments, and show the standard error of the mean.
- 4.1 Heuristic Representations and Pre-Trained Transformer Embeddings: Encoder effects are compared with Morgan fingerprints by molecular library and as effect ratios averaged across all libraries.The comparisons use 20,000 experiments with batch size 1,000; means and standard deviations come from 8,000 posterior samples, and violins show 95% highest-density intervals.
- 4.2 Transformer Pooling Strategies: MolFormer pooling strategies are assessed through optimization trajectories over 20,000 experiments using a Laplace Neural Network surrogate model.Trajectories average 20 independent campaigns with batch size 1,000 experiments and shaded standard errors of the mean.
- 4.2 Transformer Pooling Strategies: Pooling strategies are compared with mean pooling using library-specific logit-scale effects and effect ratios averaged across libraries.Virtual-screening evaluations use 20,000 experiments with batch size 1,000; posterior summaries use 8,000 samples and effect-ratio plots show 95% highest-density intervals.
- ChemBERTa: ChemBERTa fine-tuning strategies are evaluated through optimization trajectories and relative-effect analyses in virtual screening.The 20,000-experiment evaluations use batch size 1,000, 20 independent campaigns, 8,000 posterior samples for logit-scale summaries, and 95% highest-density intervals for effect ratios.
- ChemBERTa: ChemBERTa fine-tuning is additionally examined in experimental discovery using 1,000-experiment campaigns with batch size 50.The supplementary figures report trajectories, library-differentiated logit-scale effects, and effect ratios averaged over libraries.
- MolFormer-XL: MolFormer-XL fine-tuning strategies are evaluated with optimization trajectories, library-specific logit-scale effects, and cross-library effect ratios in virtual screening.These evaluations use a 20,000-experiment budget, batch size 1,000, 20 independent campaigns, 8,000 posterior samples, and 95% highest-density intervals.