Source-linked AI summary

Exploring Chemical Compound Space with Quantum-Based Machine Learning

O. Anatole von Lilienfeld, Klaus-Robert Müller, Alexandre Tkatchenko

arXiv:1911.10084v2physics.chem-phphysics.comp-ph

TL;DR

Rational compound design requires understanding and rapidly evaluating properties across the vast chemical compound space, while conventional statistical and ML approaches can lack transferability. This perspective surveys QM-based ML methods that combine physical theory, comprehensive QM/SM datasets, and chemically informed learning; it reports rapid, accurate out-of-sample prediction and a 40-fold reduction in QM9 errors, while identifying limitations in complex electronic properties and dataset representativeness.

  • Problem

    Chemical compound space is vast, and accurate QM and statistical-mechanics simulations are computationally demanding, while conventional statistical and ML approaches may not transfer beyond their application domains.

  • Method

    The perspective synthesizes QM/SM-based ML approaches that combine rigorous physical theories, comprehensive datasets, and representations incorporating physical and chemical knowledge.

  • Results

    QML models enable rapid predictions for out-of-sample systems and converged predictive power, with QM9 errors decreasing 40-fold from 8 kcal/mol in 2012 to 0.2 kcal/mol in 2018.

  • Takeaways & Limitations

    QML makes studying ensembles of compounds and recovering, discovering, and elucidating structure–property relationships feasible and valuable goals.

  • Takeaways & Limitations

    Transferable accurate QML models for electron densities, molecular orbitals, and solid-state band structures remain challenging, and training-set selection bias can hinder assessment of broader representativeness.

Abstract

from arXiv · show

Rational design of compounds with specific properties requires conceptual understanding and fast evaluation of molecular properties throughout chemical compound space (CCS) -- the huge set of all potentially stable molecules. Recent advances in combining quantum mechanical (QM) calculations with machine learning (ML) provide powerful tools for exploring wide swaths of CCS. We present our perspective on this exciting and quickly developing field by discussing key advances in the development and applications of QM-based ML methods to diverse compounds and properties and outlining the challenges ahead. We argue that significant progress in the exploration and understanding of CCS can be made through a systematic combination of rigorous physical theories, comprehensive synthetic datasets of microscopic and macroscopic properties, and modern ML methods that account for physical and chemical knowledge.

INTRODUCTION

Chemical compound space comprises feasible metastable atomic configurations, but its enormous scope makes first-principles evaluation computationally demanding. The perspective examines QML as a route to faster, physics-informed exploration by combining QM/SM theory, comprehensive data, and modern machine learning.

  • Quantum mechanics calculates microscopic properties, while statistical mechanics samples energy surfaces to obtain macroscopic properties.
  • Chemical compound space contains all feasible metastable atomic configurations resulting from solving Schrödinger’s equation.
  • Statistical and ML approaches can efficiently search CCS but often lack transferability beyond their training domains because they lack underlying physical principles.
  • Recent progress is supported by mature electronic-structure methods, high-performance computing, and adaptations of statistical mechanics and statistical learning techniques.
  • QML combines modern statistical learning with QM and SM to predict electronic and atomistic properties and processes from reference simulations.
  • The perspective emphasizes systematic integration of rigorous physical theories, comprehensive QM and SM datasets, and ML methods incorporating physical and chemical knowledge.

BOX1: EXPLANATION OF QML TERMS

The terminology box defines the representations, learning procedures, electronic-structure methods, and model families used to describe QML approaches. It also connects model evaluation to generalization, scalability, and computational accuracy.

  • Chemical Compound Space (CCS) is the set of feasible metastable atomic configurations obtained by solving the Schrödinger equation for interacting electrons and nuclei.
  • A representation encodes atomic structure and relations while remaining unique and invariant to atom indexing, translations, and rotations.
  • Supervised learning uses labeled samples for classification or regression, whereas unsupervised learning uses no label information.
  • Kernel-ridge regression provides nonlinear regression that generalizes with limited data but scales poorly for larger datasets, whereas deep neural networks scale to complex relations in large datasets.
  • Cross validation assesses generalization to unseen data and helps avoid overfitting, while learning curves measure performance as training samples increase.
  • DFT is an efficient electronic-structure method whose practical approximations balance accuracy and computational cost, and it underlies many QML datasets.
  • CCSD(T) is a high-accuracy quantum-chemistry method that typically reaches approximately 1 kcal/mol chemical accuracy for suitable molecular atomization energies.

GOALS AND ADVANCES OF QML

QML combines quantum-mechanical foundations, statistical learning, and increasingly comprehensive QM datasets to model and explore chemical compound space. Recent advances improve accuracy, transferability, scalability, and dynamical simulation, while universal force fields across configurational and compositional space remain an open challenge.

  • QML aims to develop reliable models with accuracy approaching high-level electronic-structure calculations, using CCSD(T) or DFT reference data.
  • QML differs from earlier cheminformatics approaches by grounding predictions of molecular and materials properties in quantum mechanics and statistical mechanics.
  • 40-fold lower errors on QM9—from 8 kcal/mol in 2012 to 0.2 kcal/mol in 2018—reflect advances in physical priors and molecular representations.The cited comparison uses the same dataset for training and validation; advanced models reached chemical accuracy of 1 kcal/mol on 134,000 QM9 molecules.
  • Effective representations must be unique, symmetry-aware, physically informed, and computationally efficient for reliable QML learning.They encode atomic structure and relations while capturing translational, rotational, and permutational invariances; neural networks can also learn implicit multiscale representations.
  • QML has enabled essentially exact QM-level molecular dynamics for small molecules and supported transferable DFT-level force fields for complex solid-state processes.Examples include conformational transitions and vibrational spectroscopy in ethanol and aspirin, plus growth simulations of tetrahedral amorphous carbon.
  • Universal QML force fields across chemical compound space remain challenging because simultaneous configurational and compositional navigation involves diverse chemistries and non-local quantum interactions.

QML BASED INSIGHTS INTO CCS

QML models extend quantum calculations across chemical compound space while revealing property relationships, structural motifs, and candidates for rational design. Applications span molecular dynamics, response properties, generative design, explainable representations, and crystal discovery.

  • Molecular dynamics: QML-enabled force fields allow quantum molecular dynamics with essentially exact CCSD(T)-level atomic forces for molecules as large as aspirin.This enables accurate thermodynamic and spectroscopic observables without compromising between accuracy and computational cost.
  • Response properties: Explicit derivative information qualitatively improves QML predictions of hydrogen-fluoride binding curves and dramatically improves electrostatic response-property predictions.The approach uses response operator theory to incorporate derivatives with respect to interatomic distance.
  • Inverse design: Generative models can optimize compounds in a latent chemical space, producing promising light-harvesting material candidates.Latent-space optimization provides an alternative to iteratively solving the forward compound-to-property problem.
  • Explainable QML: Analyzing learned QML representations can expose molecular-property relationships and support explainable interpretations of model predictions.DTNNs trained on quantum-mechanical molecular energies learn a finite-data approximation to the mapping between molecular structures and quantum properties.
  • Materials discovery: Nearly one hundred novel Elpasolite crystals were identified after ranking estimated formation energies for ∼2 million candidate crystals by thermodynamic stability.The same analysis found NFAl2Ca6, an exotic crystal containing aluminum with an unusual negative oxidation state.
  • Broader implications: Collectively, these examples show that QML can extract statistical insights and new knowledge about quantum-property relationships unavailable through conventional quantum calculations.The perspective frames QML as a step toward inverse design across multiple scientific domains.
  • Property relationships: Weak correlations among many molecular properties suggest that apparently conflicting design goals, including low energy, high polarizability, and high HOMO-LUMO gap, may coexist.The analysis used pairwise correlations for roughly 7k molecules in the QM7 dataset.

CHALLENGES AND OUTLOOK

Despite rapid progress, QML still faces unresolved challenges that motivate continued development of the field.

  • Challenges and outlook: The perspective identifies many challenges that remain after tremendous progress in QML.It presents these as pressing directions for future work.

Towards Big Data in CCS

QML requires larger, more comprehensive CCS datasets while remaining efficient, uncertainty-aware, explainable, and capable of active sampling. Iterative optimization can update models with new quantum-mechanical calculations until candidate generation converges.

  • Data resources: A lack of large comprehensive datasets is an important limitation for QML development.Existing datasets have supported method development and testing but may encourage overfitting to benchmarks.
  • Iterative optimization: QML application concepts iteratively combine ab initio calculations, model updates, and multiobjective candidate generation until convergence.The workflow begins with representative target compounds and spans structural and compositional degrees of freedom.
  • Data resources: High-quality large-scale resources should cover both compositional and configurational aspects of chemical compound space.The need follows from the combinatorial scaling of CCS and the limitations of current benchmark datasets.
  • Data efficiency: QML models should use physics-based priors and invariance information to learn reliably from small datasets without sacrificing robustness or accuracy.Efficient learning remains necessary because CCS scales combinatorially.
  • Model capabilities: Future models should quantify predictive uncertainty, use active learning to improve CCS sampling, and explain their predictions.The passage notes that active learning induces non-stationarity in learning.
  • Adaptive exploration: On-the-fly fragment selection and QML training represent an initial step toward more adaptive exploration of chemical compound space.The approach is presented as an early direction for improving sampling and model construction.

Learning Complex Electronic Properties

QML still faces important limitations in predicting complex electronic properties and in determining how broadly its models remain reliable across chemical compound space.

  • Current QML limitations include predicting molecular-orbital eigenvalues and excited-state properties such as excitation energies.
  • Transferable and accurate QML models of electron densities, molecular orbitals, and band structures remain challenging.
  • Training-set selection bias hampers rigorous assessment of whether datasets represent broader chemical spaces.Stability and property distributions are typically unknown in chemical compound space.
  • Open conceptual challenges include identifying effective dimensionality, controlling prediction errors across chemical space, and relating learning efficiency to data dimensionality.

Multiscale QML models

Multiscale QML combines data from different theory levels so models learn corrections between approximations, while the broader goal is universal exploration beyond restricted chemical spaces.

  • Multiscale QML models: Δ-learning combines abundant low-cost approximate data with fewer high-accuracy points, reducing the data required to learn differences between theory levels.This allocates model complexity to the most difficult aspects of the prediction problem.
  • Multiscale QML models: The field aims to extend successful structure–property modeling from restricted chemical spaces toward global and universal exploration of chemical compound space.
  • Multiscale QML models: Available QML tools have reached a maturity level considered helpful to a wider community of researchers and practitioners.

Towards molecular design with QML

The path toward molecular design with QML involves developing tools for more observables, automated experimental design, and reaction discovery beyond historically represented chemistry.

  • More observables: Predicting statistical-mechanics observables such as free-energy profiles of rare events remains an outstanding challenge.Solvent effects must also be included to compare QML predictions with liquid-chemistry results.
  • Experimental design: QML can support automated experiments and experimental-design decisions, including multi-property design through latent spaces, computational alchemy, and optimizers.
  • Reaction design using QML: Literature-based reaction models are biased toward historically known chemistries and may not represent new combinations or reaction conditions.
  • Reaction design using QML: Novel reaction profiles require first-principles QM and SM simulations that account for electronic and atomic movements, reaction energies, and barriers.

CONCLUSIONS

QML models can provide rapid, convergent predictions across chemical compound space and make structure–property analysis more feasible, but broad validity, education, and full control remain unresolved challenges.

  • QML models enable rapid predictions for out-of-sample systems and converged predictive power with sufficiently large training sets.Learning-curve convergence provides evidence for the latter capability.
  • The reduced cost of query tasks makes recovery, rejection, discovery, and elucidation of structure–property relationships feasible across compound ensembles.
  • Defining locality and model-validity boundaries, quantifying uncertainty, and measuring diversity remain important conceptual challenges across diverse material and molecular systems.
  • The interdisciplinary field needs tightly integrated chemistry, physics, and computer-science curricula that conventional departments do not currently provide.
  • Routine discovery and design of molecules and materials with desired properties remains an unrealized long-term goal.

Appendix: Software

The appendix points readers to software packages implementing the quantum machine-learning methods discussed in the paper.

  • General QML code is available in reference, while SchNetpack supports implementations of deep-learning methods such as DTNN and SchNet.
Loading 1911.10084v2…