Source-linked AI summary

Big Data meets Quantum Chemistry Approximations: The $Δ$-Machine Learning Approach

Raghunathan Ramakrishnan, Pavlo O. Dral, Matthias Rupp, O. Anatole von Lilienfeld

arXiv:1503.04987v1physics.chem-ph

TL;DR

The computational cost of quantum chemistry limits exhaustive studies of molecular chemical space. This paper combines inexpensive approximate quantum methods with machine-learning corrections, achieving high accuracy on larger molecular sets, including over 100,000 organic molecules.

  • Problem

    The computational demands of quantum chemistry make routine investigation of larger subsets of chemical space difficult, despite the importance of screening possible molecules.

  • Method

    The Δ-ML approach combines fast legacy quantum-chemical approximations with machine-learning models trained on accurate reference results to predict corrections for new molecules.

  • Results

    Semi-empirical quantum-chemistry error decreased from 7.2 kcal/mol to approximately 5 kcal/mol or 3 kcal/mol for over 100,000 organic molecules using less than 1% or 10% for training, respectively.

  • Takeaways & Limitations

    The results support Δ-ML as a strategy for augmenting legacy quantum chemistry with modern machine learning while reducing computational burden.

Abstract

from arXiv · show

Chemically accurate and comprehensive studies of the virtual space of all possible molecules are severely limited by the computational cost of quantum chemistry. We introduce a composite strategy that adds machine learning corrections to computationally inexpensive approximate legacy quantum methods. After training, highly accurate predictions of enthalpies, free energies, entropies, and electron correlation energies are possible, for significantly larger molecular sets than used for training. For thermochemical properties of up to 16k constitutional isomers of C$_7$H$_{10}$O$_2$ we present numerical evidence that chemical accuracy can be reached. We also predict electron correlation energy in post Hartree-Fock methods, at the computational cost of Hartree-Fock, and we establish a qualitative relationship between molecular entropy and electron correlation. The transferability of our approach is demonstrated, using semi-empirical quantum chemistry and machine learning models trained on 1 and 10\% of 134k organic molecules, to reproduce enthalpies of all remaining molecules at density functional theory level of accuracy.

I. INTRODUCTION

The paper addresses the cost of accurate quantum chemistry across large chemical spaces by learning corrections to efficient baseline methods. The Δ-ML approach targets property, theory-level, and geometry differences while supporting out-of-sample prediction across chemically diverse molecules.

  • Motivation: Chemical-space screening is constrained because exhaustive investigation of potentially interesting molecules and accurate quantum-chemical calculations are computationally demanding.Chemical accuracy is approximately 1 kcal/mol and can matter for reaction rates, structure–property relationships, and molecular design.
  • Motivation: Approximate methods such as PM7, Hartree–Fock, and DFT capture much of the relevant physics, leaving a smaller energy contribution for correction.For H2O, Hartree–Fock approximates the experimental ionization potential by more than 90%.
  • Δ-ML approach: The Δ-model estimates a target property from a baseline value plus a machine-learned correction between baseline and target levels.The correction can account for changes in observable, theory level, and molecular geometry.
  • Δ-ML approach: The ML correction uses kernel ridge regression over similarities between sorted Coulomb-matrix representations of query and training molecules.The representation is invariant to translation, rotation, and atom indexing, except among enantiomers.
  • Evidence: With 1k training molecules, Δ-ML reduced GW HOMO MAE from 0.78 to 0.23 eV for ZINDO and from more than 2 to less than 0.1 eV for PBE0.GW LUMO MAE likewise fell from 0.91 to 0.16 eV for ZINDO and from 1.3 to 0.13 eV for PBE0.
  • Evidence: The study applies this strategy to chemically diverse C7H10O2 isomers with dense, near-degenerate atomization-energy distributions.The dataset includes up to approximately 100 molecules per kcal/mol of atomization enthalpy and near-degenerate molecules within approximately 0.01 kcal/mol.

A. Chemically accurate prediction of covalent bonding

For 6k constitutional isomers of C7H10O2, Δ-ML errors decrease systematically with training-set size and can reach chemical accuracy for high-level atomization energies. The approach outperforms baseline-free ML and fixed reparameterizations, with descriptor choice also affecting performance.

  • Accuracy scaling: The out-of-sample MAE rapidly decreases with training-set size and maintains a constant decay rate beyond 1k training molecules.The figure compares atomization-energy predictions against G4MP2 references for different baselines.
  • Accuracy scaling: Less than 1k training molecules brought both PBE and B3LYP ΔG4MP2 models to chemical accuracy below 1 kcal/mol.The B3LYP-based model reached approximately 0.4 kcal/mol with 5k training molecules.
  • Accuracy scaling: 2k training molecules gave PM7-based ΔG4MP2 accuracy similar to pure B3LYP, while 5k reduced the error below 2 kcal/mol.The PM7 baseline is substantially faster but more approximate than the DFT baselines.
  • Model comparisons: All Δ-ML models outperform direct ML trained on absolute G4MP2 atomization energies without a baseline.This comparison is reported for the C7H10O2 isomer set.
  • Model comparisons: Reparameterization improves the initial offset but its advantage vanishes beyond 1k training molecules, and OM2 shows a similar pattern.The results indicate limitations of fixed functional forms with globally optimized parameters.
  • Model comparisons: Replacing the Coulomb-matrix representation with a bag-of-bonds descriptor produced similar decay rates and reached chemical accuracy with 5k training molecules.The bag-of-bonds model performed better than the Coulomb-matrix-based Δ-model in this comparison.

B. Chemically accurate thermochemistry

The thermochemistry study extends Δ-ML from potential energies to thermodynamic properties at 298.15 K. Models learn theory, geometry, and thermal corrections while keeping prediction cost dominated by the inexpensive baseline evaluation.

  • Thermochemical properties: Δ-ML predicts internal energies, enthalpies, free energies, and entropies of atomization for 6k C7H10O2 isomers at G4MP2 level.The models use potential energy of atomization as the baseline property and learn the remaining thermodynamic effects.
  • Thermochemical properties: The 1k-ΔG4MP2 PM7-ML model reduced PM7 prediction error and standard deviation by more than approximately 50%.The model converged to near chemical accuracy of 1.7 kcal/mol with a 5k training set.
  • Thermochemical properties: Internal energies, free energies, and entropies showed nearly identical convergence and baseline trends to the enthalpy predictions.This indicates consistent training-set scaling across the reported thermodynamic properties.
  • Computational cost: Out-of-sample computational effort is dominated by baseline evaluations rather than the learned correction.The approach avoids calculating corresponding partition functions with more accurate theories, which can be prohibitively expensive.

C. Electron correlation

The Δ-ML approach models electron-correlation corrections from inexpensive Hartree–Fock or post-Hartree–Fock baselines. On 6k C7H10O2 isomers, training corrections substantially reduce errors, including to below chemical accuracy.

  • Computational motivation: MP2 scales as N_e^5, while CCSD(T) scales as N_e^7, motivating machine-learning corrections for post-Hartree–Fock correlation energies.Electron correlation is defined as the difference between converged-basis HF energy and the corresponding non-relativistic exact result.
  • Evaluation: For 6k C7H10O2 isomers, Δ-ML corrections were evaluated for combinations of HF, MP2, CCSD, and CCSD(T), after removing systematic shifts.The comparisons included baseline and corrected atomization-energy predictions.
  • Accuracy: The ΔCCSD(T) correction from HF reduced the correlation-energy MAE from ∼2.9 to less than 1 kcal/mol.This spans the least and most expensive methods considered in the comparison.
  • Accuracy: The 1k-ΔMP2 correction from HF had a larger MAE than 1k-ΔCCSD from HF, suggesting that MP2 is a more complex function in chemical space.The comparison concerns prediction of the respective higher-level targets, despite MP2 being more approximate than CCSD.
  • Scaling with training size: For larger training sets, correlation-energy errors decayed similarly to errors from DFT baseline models for thermodynamic properties.The least approximate baseline, 1k-ΔCCSD(T) from CCSD, yielded a chemically nearly negligible MAE of ∼0.1 kcal/mol.

D. Applicability: Diastereomers of C7H10O2

The authors tested Δ-ML transferability by screening nearly 10k C7H10O2 diastereomers using models trained on 1k constitutional isomers. The DFT-based correction reproduced G4MP2 enthalpies within chemical accuracy for validation cases and identified the same global minimum.

  • Screening: A 1k-ΔG4MP2 model trained on constitutional isomers screened all 9868 unique, stable C7H10O2 diastereomers.Validation used G4MP2 enthalpies for a randomly selected 3k diastereomers.
  • Stable-isomer identification: The model predicted 6-oxabicyclooctan-7-one as the most stable isomer with H = -1933.5 kcal/mol, matching a validating G4MP2 calculation.Its ten closest isomers occupied a narrow 9 kcal/mol enthalpy window.
  • Reaction energetics: For the ten closest products, predicted isomerization enthalpy errors never exceeded 1 kcal/mol, with a maximal error of 0.6 kcal/mol for product 10.The Δ-ML predictions agreed with G4MP2 results calculated a posteriori.
  • Reaction energetics: B3LYP alone found the same global minimum but produced substantially deviating reaction enthalpies, including a spectacular failure for isomer 8.The corrected model retained DFT computational cost while accounting for subtle baseline errors in competitive bonding.

E. Interpretation of the ∆-Model:

The Δ-model correction can represent correlation energy and, through thermochemical differences, atomization entropy. Applied to C7H10O2 isomers and diastereomers, it revealed a qualitative but non-universal relationship between entropy and electron correlation.

  • Model interpretation: A ΔCCSD(T)-HF atomization-energy model can be interpreted as a machine-learning model of atomization correlation energy.This interpretation follows because the HF baseline is corrected toward CCSD(T).
  • Model interpretation: Subtracting free-energy and enthalpy atomization models cancels the baseline energy and, after division by T, yields atomization entropy.The relation is expressed through the thermodynamic difference between H and G.
  • Screening: Using 1k training isomers and PM7 geometries, the authors trained models for G4MP2 atomization entropy and CCSD(T) electron-correlation energy.The models were reapplied to screen approximately 10k diastereomers for extreme entropy and correlation-energy differences.
  • Molecular interpretation: The most compact molecule with few degrees of freedom exhibited maximal correlation, whereas the most elongated molecule had the least correlation energy.For entropy, the compact cage-like isomer had the lowest entropy and the flexible elongated molecule the largest.
  • Relationship between properties: Across parent isomers and diastereomers, entropy and correlation energy showed a qualitative, hardly quantitative interdependence associated with molecular compactness or extension.The authors suggest this as a molecular analogue of phonon–electron coupling phenomena in solids.
  • Scope: The entropy–correlation trend breaks down for organic molecules of very different sizes in GDB-9 and may require normalization by atom or electron count.Thus, the observed relationship is not established as general across molecular sizes.
  • Computational cost: For 10k diasteromers, PM7 geometry relaxations required ∼1 CPU day, while the ML correction to G4MP2-S and CCSD(T)-Ec was instantaneous.The corresponding high-throughput ab initio calculations were estimated to require ∼20 CPU years.

F. Thermochemistry for 134 kilo organic molecules

The authors applied Δ-ML to 134k organic molecules to approximate B3LYP enthalpies from a PM7 baseline. Training on 1k or 10k molecules reduced errors across the remaining chemical space while retaining PM7-scale screening cost.

  • Motivation and approach: Hierarchical screening uses inexpensive methods to filter compounds before applying more accurate but costlier quantum chemistry, but DFT is too expensive for hundreds of thousands of molecules.The Δ-approach targets DFT-quality filtering at semi-empirical computational cost.
  • Baseline: PM7 atomization enthalpies deviated from B3LYP by 7.2 kcal/mol on average before machine-learning correction.This establishes the baseline error that Δ-ML seeks to correct.
  • Accuracy: A 1k-ΔB3LYP-PM7 model predicted 133k out-of-sample B3LYP enthalpies with an MAE of 4.8 kcal/mol, while 10k training reduced the MAE to 3.0 kcal/mol on 124k molecules.The errors were measured on molecules excluded from the respective training sets.
  • Accuracy: The 10k-ΔB3LYP-PM7 model was reported as comparable to GGA or hybrid DFT while retaining PM7 computational cost.The ML correction removed systematic underestimation and skew already with 1k training molecules, with further contraction at 10k.
  • Geometry dependence: The correction reduced geometry-dependent PM7 deviations from approximately 20 kcal/mol across much of the molecular-shape space into a 5 kcal/mol error window.A few 20 kcal/mol outliers persisted near the rod–disk edge.
  • Computational cost: Screening all 134k molecules took less than 2 CPU weeks with ΔB3LYP-PM7, compared with an estimated 15 CPU years for direct B3LYP screening.A single corrected evaluation required no more than 10 seconds for the largest GDB-9 molecule.

G. Conclusions

The Δ-ML approach combines approximate quantum chemistry with machine learning to predict high-level properties for out-of-sample molecules at substantially lower computational cost. It identifies competitive isomers, reveals a qualitative entropy–correlation relationship, and transfers across large organic-molecule sets.

  • The Δ-ML model reaches high-level quantum-chemistry accuracy for out-of-sample molecules while computational cost is dominated by the fast baseline method.The approach augments PM7, HF, or DFT with machine-learning corrections trained on accurate reference results.
  • For 10k C7H10O2 diastereomers, the method identifies the ten most competitive reaction isomers of the most stable isomer.
  • A qualitative dependency between entropy and atomization correlation energy suggests a molecular equivalent of electron-phonon coupling.
  • Using less than 1% and 10% training subsets of 134k organic molecules, Δ-ML reduces semi-empirical error from 7.2 kcal/mol to approximately 5 and 3 kcal/mol, respectively.The approximate 5 kcal/mol target corresponds to generalized-gradient-approximated theory, while approximately 3 kcal/mol corresponds to hybrid density functional theory.
  • The approach may extend predictive improvements to heat capacities, non-adiabatic corrections, reaction barriers, optical properties, atomic forces, semi-empirical parameters, and electronic excitations.

A. Molecular datasets

The study uses four organic-molecule datasets spanning preliminary molecular properties, a 134k-molecule GDB collection, accurate C7H10O2 isomer calculations, and additional data for demonstrating versatility.

  • The first dataset contains 7,211 organic molecules with HOMO/LUMO eigenvalues and molecular polarizabilities at multiple theory levels.It is used for preliminary testing of the approach.
  • The second dataset contains 133,885 GDB molecules with up to nine heavy atoms, with PM7 and B3LYP thermochemical properties including enthalpies and entropies of atomization.
  • A subset of 6,095 C7H10O2 constitutional isomers has thermochemical properties calculated at a more sophisticated level considered to provide chemical accuracy of approximately 1 kcal/mol.
  • The 133,885-molecule collection excludes cations, anions, and molecules containing S, Br, Cl, or I.
Loading 1503.04987v1…