Source-linked AI summary

Machine Learning of Molecular Electronic Properties in Chemical Compound Space

Grégoire Montavon, Matthias Rupp, Vivekanand Gobre, Alvaro Vazquez-Mayagoitia, Katja Hansen, Alexandre Tkatchenko, Klaus-Robert Müller, O. Anatole von Lilienfeld

arXiv:1305.7074v1physics.chem-ph

TL;DR

Electronic-structure calculations produce vast datasets, but extracting predictive structure–property relationships remains difficult. The paper develops a deep multi-task neural network that uses Coulomb-matrix representations to predict multiple molecular properties simultaneously. For small organic molecules, the model reaches reference-level accuracy and enables virtually instantaneous predictions.

  • Problem

    Extracting predictive structure–property relationships from rapidly growing electronic-structure datasets remains challenging.

  • Method

    A deep multi-task neural network maps Coulomb-matrix molecular representations to multiple electronic properties simultaneously.

  • Results

    Errors for all properties are in the single-digit percentage range of their means, with accuracy comparable to the corresponding reference methods after training on 5000 molecules.

  • Takeaways & Limitations

    The Quantum Machine provides simultaneous, virtually instantaneous predictions and supports exploration of chemical compound space for designing novel and improved compounds.

Abstract

from arXiv · show

The combination of modern scientific computing with electronic structure theory can lead to an unprecedented amount of data amenable to intelligent data analysis for the identification of meaningful, novel, and predictive structure-property relationships. Such relationships enable high-throughput screening for relevant properties in an exponentially growing pool of virtual compounds that are synthetically accessible. Here, we present a machine learning (ML) model, trained on a data base of \textit{ab initio} calculation results for thousands of organic molecules, that simultaneously predicts multiple electronic ground- and excited-state properties. The properties include atomization energy, polarizability, frontier orbital eigenvalues, ionization potential, electron affinity, and excitation energies. The ML model is based on a deep multi-task artificial neural network, exploiting underlying correlations between various molecular properties. The input is identical to \emph{ab initio} methods, \emph{i.e.} nuclear charges and Cartesian coordinates of all atoms. For small organic molecules the accuracy of such a "Quantum Machine" is similar, and sometimes superior, to modern quantum-chemical methods---at negligible computational cost.

I. INTRODUCTION

The paper addresses how to extract predictive structure–property relationships from rapidly growing electronic-structure datasets. It presents a first-principles machine-learning approach that maps molecular nuclear charges and positions to multiple electronic properties.

  • Electronic-structure calculations can generate data for millions of virtual compounds, but extracting predictive structure–property relationships remains challenging.
  • Existing QSPR approaches typically regress descriptor variables against properties, requiring heuristic descriptor design for each property and chemical class.
  • The proposed approach uses nuclear charges and atomic positions—the variables entering the electronic Hamiltonian—to infer molecular properties.
  • A deep multi-task neural network simultaneously predicts diverse electronic properties while exploiting correlations among properties and theory levels.
  • The training database contains nearly 10^5 entries for more than 7000 small organic molecules, supporting out-of-sample prediction.

II. METHODS

The study uses a controlled database of 7211 small organic molecules and represents each molecule with randomized Coulomb matrices. The database combines quantum-chemistry calculations, molecular geometries, and a multi-property visualization of training and testing data.

  • The test bed contains 7211 small organic molecules with up to seven non-hydrogen atoms from C, N, O, S, and Cl, saturated with hydrogens.
  • Molecular geometries are generated from SMILES with the universal force field and relaxed using PBE Kohn–Sham DFT.
  • The molecular representation is derived solely from stoichiometry and configuration, without pre-conceived chemical or quantum-mechanical descriptors.
  • The randomized Coulomb matrix encodes inverse atom distances, is unique up to identical molecules or enantiomers, and is invariant to translation and rotation.
  • Random Coulomb matrices represent different atom indexings of the same molecule as a probability distribution, enforcing atom-indexing invariance.
  • Figure 1 summarizes quantum-chemistry results for 14 properties across 7211 molecules and illustrates the model’s input and latent-property decoding.

C. Molecular electronic properties (output)

The model targets electronic ground- and excited-state properties computed for molecules at their PBE geometry minima. Multiple reference methods are used to assess the influence of theory level while balancing computational cost and predictive accuracy.

  • The outputs include atomization energy, static polarizability, HOMO and LUMO eigenvalues, ionization potential, electron affinity, and excitation-spectrum quantities.
  • PBE0 is used for atomization energies and frontier orbital eigenvalues, while SCS and PBE0 provide alternative polarizability references.
  • ZINDO supplies electron affinity, ionization potential, excitation energies, and maximal absorption intensity for the optical-spectrum calculations.
  • GW evaluates frontier orbital eigenvalues as a quasiparticle many-body perturbation method for electron addition and removal processes.
  • The selected theory levels are described as a compromise between computational cost and predictive accuracy, while machine learning can in principle use any approximation level.

D. Training the model

The model is a deep multi-task neural network trained on molecule–property pairs. It maps Coulomb matrices to all 14 molecular properties simultaneously, using shared constraints across outputs.

  • The neural network is trained on molecule–property pairs to map Coulomb matrices to all 14 properties simultaneously.
  • Multi-task learning constrains the model to fit multiple properties at once, reducing the set of models that can satisfy the data simultaneously.
  • The approach relies on neural networks’ functional-approximation capabilities and an underlying noise-free Schrödinger equation.

III. RESULTS AND DISCUSSION

The database reveals both expected and less obvious relationships among molecular electronic properties, including correlations across methods and hidden associations that the neural network can exploit.

  • Ionization potential correlates with HOMO eigenvalues, polarizability with stability, and electron affinity with first excitation energy.
  • PBE0 and SCS polarizabilities are strongly correlated across the molecular database.
  • Subtracting 1.5 eV from PBE0 HOMO values obtains GW HOMO eigenvalues to a very decent degree.
  • The neural network extracts hidden correlations even where scatter plots show little visual correlation, including atomization energy versus HOMO.

B. Accuracy vs. training set size

Prediction error decreases as the training set grows, with atomization energy showing especially steep improvement; 5000 examples already reach reference-method precision for almost all properties.

  • A ten-fold increase in training molecules from 500 to 5000 reduces atomization-energy MAE by 70%, from 0.55 to 0.16 eV.
  • For all investigated properties, MAE improvement with increasing training size suggests further reductions could result from adding molecules.
  • The reference method’s precision is reached for almost all properties using 5000 training examples, limiting the value of additional examples in this study.
  • The error decay follows ∝1/N only for atomization energy; other properties decay more slowly.
  • Figure 2 reports MAE and error bars for atomization energy, polarizability, HOMO, LUMO, and first excitation energy on logarithmic training-size axes.

C. Final ML model

The final multi-task model predicts 14 properties for 2211 out-of-sample molecules with errors comparable to the reference methods, while requiring milliseconds rather than CPU hours per molecule.

  • 2211 out-of-sample molecules receive all 14 quantum-chemical property predictions after cross-validated training on 5000 molecules.
  • Errors for all properties are in the single-digit percentage range of the mean property, with accuracy similar to the reference methods except for the most intense absorption.
  • The model’s predictive power is rationalized by deep representation learning, randomized Coulomb matrices, and multi-task exploitation of property correlations.
  • Evaluating all 14 properties for an out-of-sample molecule takes milliseconds with ML instead of several CPU hours using the reference methods.
  • Transferability is limited: accurate predictions should not be expected for compounds that do not resemble the training set.

IV. CONCLUSION

The Quantum Machine uses deep multi-task neural networks to predict multiple molecular electronic properties, offering rapid predictions while remaining limited to interpolation within represented chemical space.

  • PCA indicates that successive neural-network layers extract representations of chemical space that better capture multiple molecular properties.
  • A single Quantum Machine execution simultaneously yields multiple properties at multiple levels of theory.
  • Increasing the training set can systematically reduce error until accuracy outperforms modern quantum chemistry methods, including hybrid density-functional theory and GW.
  • The model makes virtually instantaneous property predictions and can learn from numerical calculations or experimental results at arbitrary reference levels.
  • The study provides numerical evidence that ML can infer predictive structure-property relationships from high-quality first-principles or experimental databases.
  • The demonstrated study covers small organic molecules with up to seven non-hydrogen atoms, and similar accuracy elsewhere may require different training-data amounts.
  • Reliable databases combined with ML are presented as a step toward computational bottom-up design of novel and improved compounds.

VI. SUPPLEMENTARY MATERIALS

The supplementary materials provide access to true properties and geometries for all compounds through the Quantum Machine website.

  • True properties and geometries for all compounds can be retrieved from www.quantum-machine.org.

Appendix A: Details on Random Coulomb Matrices

Random Coulomb matrices represent the same molecule under varied atom indexings by perturbing row norms before applying a shared row-and-column permutation.

  • Random Coulomb matrices define a probability distribution over matrices representing different atom indexings of the same molecule.
  • The procedure computes each row norm, adds zero-mean unit-variance noise, and sorts the resulting values to determine a permutation.
  • The same permutation is applied to Coulomb-matrix rows and columns, preserving the molecule’s representation while varying atom ordering.

Appendix B: Details on Training the Neural Network

For a new molecule, the model converts nuclear charges and coordinates into a Coulomb-matrix representation, binarizes it, propagates it through the neural network, and rescales predicted properties.

  • Prediction begins by entering Cartesian coordinates and nuclear charges, forming a Coulomb matrix, and binarizing its representation.
  • Binarization distributes continuous Coulomb-repulsion information across many dimensions with low individual information content.
  • With granularity θ = 1, a 23 × 23 Coulomb matrix becomes a quasi-binary tensor with approximately 2000 non-constant dimensions.
  • With granularity 0.25, a 14-property vector becomes a matrix with approximately 1000 non-constant components.
  • The four-layer network uses 2000, 800, 800, and 1000 nodes, applies learned linear transformations with sigmoid nonlinearities, and minimizes mean absolute error using stochastic gradient descent.
Loading 1305.7074v1…