Source-linked AI summary
Machine Learning Unifies the Modelling of Materials and Molecules
Albert P. Bartok, Sandip De, Carl Poelking, Noam Bernstein, James Kermode, Gabor Csanyi, Michele Ceriotti
TL;DR
Accurately modelling molecular and condensed-phase stability remains difficult because material and chemical modelling strategies have developed separately. The paper presents SOAP-GAP, a unified local-environment and Bayesian learning framework that succeeds across materials, molecules and biological systems, including accuracy below 0.2 kcal/mol with ntrain < 20,000.
Problem
Material and chemical modelling strategies have remained disconnected despite the fundamental importance of calculating molecular and condensed-phase energies.
Method
SOAP-GAP uses similarities between local atomic environments with atom-based kernel ridge regression to decompose total energy into atomic contributions.
Results
The framework shows consistent success across materials, molecules and biological systems, with ntrain < 20,000 allowing accuracy of less than 0.2 kcal/mol.
Takeaways & Limitations
The framework provides a unified and systematic way to model atomic-scale properties across materials, molecules and biological systems.
Takeaways & Limitations
The model relies on a training dataset produced through an ad hoc process.
Abstract
from arXiv · showhide
Determining the stability of molecules and condensed phases is the cornerstone of atomistic modelling, underpinning our understanding of chemical and materials properties and transformations. Here we show that a machine learning model, based on a local description of chemical environments and Bayesian statistical learning, provides a unified framework to predict atomic-scale properties. It captures the quantum mechanical effects governing the complex surface reconstructions of silicon, predicts the stability of different classes of molecules with chemical accuracy, and distinguishes active and inactive protein ligands with more than 99% reliability. The universality and the systematic nature of our framework provides new insight into the potential energy surface of materials and molecules.
1 Introduction
Accurate stability prediction for molecules and condensed phases is essential but computationally demanding. The paper presents a unified machine-learning framework combining Gaussian process regression with a general local descriptor to achieve predictive accuracy across matter and molecules.
- Motivation: Stability prediction requires energies accurate to approximately 0.5 kcal/mol at room temperature, a small fraction of chemical-bond energies.The cited passage gives ∼0.5 kcal/mol as the thermal energy at room temperature and up to ∼230 kcal/mol for an N2 bond.
- Limitations: Electronic-structure methods provide the required accuracy, but their computational cost limits routine applications to dozens of atoms for CC and hundreds for DFT.Exact electronic-structure solutions are described as prohibitively costly, while CC and DFT remain significant-cost models.
- Limitations: Force fields reduce cost but have generally remained qualitative, while pursuing low-cost accuracy has sacrificed generality across different systems.The introduction contrasts empirical potentials with the disconnected treatment of materials and chemistry in prior approaches.
- Contribution: Combining Gaussian process regression with a general, systematic local descriptor reunites modelling of hard matter and molecules with predictive accuracy.The framework is presented as a foundation for a universal reactive force field that can recover Schrödinger-equation accuracy at negligible cost.
- Scope: The framework also accurately classifies active and inactive protein ligands, indicating extension to more complex non-local properties.This result provides evidence that a local model can capture properties extending beyond directly local interactions.
2 Results
SOAP-GAP models capture quantum-mechanically governed silicon reconstructions and achieve chemical-accuracy predictions for molecular energies and receptor–ligand binding. Bayesian error estimates and structured sampling further improve training efficiency and reveal inconsistencies in molecular data.
- Silicon surfaces: SOAP-GAP quantitatively describes silicon dimer tilts and correctly orders Si(111) reconstructions without explicitly considering quantum mechanics.The model was trained on short ab initio molecular-dynamics trajectories of small unit cells and also describes bulk, defect, pressure, temperature, and transition-state properties.
- Silicon surfaces: Predicted error σ∗ identifies surface-adatom configurations for inclusion, bringing the model closer to accurate reconstruction predictions.The adatoms had the highest predicted error, motivating targeted additions of small surface unit cells containing adatoms.
- Molecular energies: CCSD(T) atomization energies achieve error below 1 kcal/mol with only 500 training points and less than 0.2 kcal/mol using 15% of GDB9.Using DFT energies as a baseline and learning ΔDFT-CC produces a five-fold test-error reduction versus learning CC energies directly.
- Molecular energies: A 0.26 kcal/mol accuracy is reached with a multi-scale kernel, while GDB9 reaches 1 kcal/mol with 5000 reference energies and 0.18 kcal/mol with 75000 structures.The same approach can reduce the reference set needed for 1 kcal/mol accuracy to fewer than 1000 farthest-point-sampled points in another dataset.
- Molecular energies: SOAP-GAP can learn from forces, enabling models for local fluctuations and chemical reactivity through non-equilibrium training configurations.This extension builds on the demonstrated silicon force field and prior elemental applications.
3 Discussion
The discussion shows how SOAP-GAP reveals locality and spatial structure in learned energy corrections while extending predictions across materials, molecules, and biological systems. It also identifies limitations of additive local models and directions for capturing longer-range and more complex interactions.
- Electronic-structure corrections: SOAP-GAP errors are similar for generalized-gradient and hybrid DFT, indicating that exact-exchange corrections are particularly short-ranged and learnable with local kernels.The additive SOAP kernel also decomposes method differences into atom-centered contributions, revealing cancellation between carbon and aliphatic-hydrogen terms and localization near chemically important groups.
- Ligand binding: Accurate ligand-binding prediction requires capturing spatial correlations between molecular features, which averaged atomic SOAP kernels fail to represent satisfactorily.Environment matching introduces non-locality by pairing environments between molecules rather than averaging them.
- Model interpretability: SOAP binding-score projections identify conserved warheads in active ligands, while blue regions suggest target locations for lead optimization.For decoys, no consistent patterns are resolved; in the A2 receptor example, positive fields localize to fragments occupying the binding pocket.
- General framework: SOAP-GAP’s consistent performance across materials, molecules, and biological systems links molecular geometry directly to stability without explicit electronic-structure or free-energy calculations.The framework’s rigorous local-environment representation also captures complex non-local properties, with further development proposed through deep learning, multiscale kernels, and active learning.
4 Materials and Methods
The method uses Gaussian process regression with symmetry-preserving SOAP descriptors of local atomic environments to model molecular and periodic-structure properties. Averaging local-environment similarities yields structure-level kernels and enables Gaussian Approximation Potentials that decompose total energies into atomic contributions.
- Gaussian process regression: Gaussian process regression uses a kernel as a similarity measure and provides both posterior predictions and an estimate of prediction error.Its regularization parameter represents expected deviations from the underlying model due to statistical or systematic errors.
- Local-environment representation: The learning task is reduced by representing molecules and solids through local atomic environments and kernels between those environments.This focuses the model on local features rather than the vast space of all possible molecules and solids.
- SOAP descriptor and kernel: The SOAP kernel compares neighbor densities within a finite cutoff, integrates over three-dimensional rotations, and respects rotational, translational, and permutation symmetries.Its associated spherical power spectrum provides a chemical descriptor that is smooth in atomic coordinates and refinable toward a complete environment description.
- Structure-level kernel: Structure-level kernels are constructed by averaging the SOAP similarity over all pairs of atomic environments in two molecules or periodic structures.For energy-per-atom fitting, choosing uniform pair weights is equivalent to summing atomic energy contributions.
- Gaussian Approximation Potential: The model learns an optimal decomposition of quantum-mechanical total energy into atomic contributions from total energies and atomic derivatives, defining a Gaussian Approximation Potential.A SOAP-GAP model specifically uses the SOAP kernel as its basis.
Supporting Materials
The supporting materials establish the equivalence between molecular structure-kernel learning and atom-based energy decomposition, then describe the SOAP-GAP silicon model and its physically motivated construction. They also report that the model achieves high accuracy in subtle surface-reconstruction energetics while noting that detailed comparisons will be published separately.
- The atom-centered GAP is equivalent to the average molecular kernel: The atom-based formulation evaluates contributions from individual atomic environments and corresponds to the conventional definition of an interatomic potential.Its kernel measures similarity between atomic environments, building on earlier GAP applications to materials.
- The atom-centered GAP is equivalent to the average molecular kernel: Molecular energies learned with kernels averaged over atomic environments are equivalent to learning an additive atom-based energy decomposition.The equivalence follows by relating molecular and atomic kernels to energy covariances under a fully additive decomposition.
- A SOAP-GAP potential for silicon: The silicon SOAP-GAP training configurations were generated with DFT molecular dynamics, after which energies, forces, and virials were recalculated using tighter convergence settings.The tighter settings included a 250 eV plane-wave cutoff and 0.03 Å−1 k-point density.
- A SOAP-GAP potential for silicon: The silicon model uses a 5 Å locality cutoff and 0.5 Å neighbor-density width, with hyperparameters chosen from physically motivated guesses rather than optimized.Prior calculations indicated that force errors below 0.1 eV/Å could be achieved with a cutoff of around 5 Å.
- A SOAP-GAP potential for silicon: The model combines a SOAP-GAP term with a pair potential fitted to Si-dimer dissociation and close-range repulsion behavior.The pair term augments the Gaussian-process model where short-bond energies are much larger than interatomic bonding attractions.
- A SOAP-GAP potential for silicon: The results demonstrate exquisite accuracy in subtle situations such as silicon surface reconstructions, while detailed comparisons with other widely used potentials are deferred.The supporting text characterizes this as going beyond previously reported accuracy near training-set configurations, but the passage is truncated before listing all claimed advances.
A. Computational details
The computational workflow combined database geometries with PM7-based optimization and higher-level quantum calculations. Molecular and oligopeptide structures were treated using specified software, functionals, basis sets, and coupled-cluster energetics.
- Molecular geometry generation: DFT geometries and energies came from the original GDB9 database, while PM7optimised geometries were generated from GDB9 SMILES strings.CORINA version 3.60 0066 constructed three-dimensional models and initial Cartesian coordinates before PM7 relaxation.
- Molecular geometry generation: The resulting molecular configurations were relaxed at the PM7 level using MOPAC versions 16.043L and 17.048L.The workflow used MOPAC for semi-empirical PM7 geometry relaxation.
- Oligopeptide and higher-level calculations: Oligopeptide relaxations used Gaussian 09 with DFT, the B3LYP functional, and the 6-31G(2df,p) basis set to match GDB9.CCSD(T) energetics of the DFT-relaxed configurations were calculated with MOLPRO version 2012.1 using the 6-311G** basis set.
B. Training set selection and error distribution
Training-set selection strongly affects error distributions in the unevenly sampled GDB9 chemical space. Farthest-point sampling broadens structural coverage and reduces RMS error, despite a marginally higher MAE than random selection.
- Error distribution: ntrain = 20, 000 training structures were selected either randomly or by FPS to compare error distributions and learning curves.Figure S2 compares the fraction of test configurations below error thresholds and reports MAE and RMS learning curves.
- Training set selection: Farthest-point sampling (FPS) promotes uniform coverage of chemical space, including sparsely represented margins and densely sampled regions.FPS greedily selects molecules farthest from existing references using a kernel-induced structural metric.
- Error distribution: FPS produces a marginally higher MAE than random selection but significantly reduces RMS error.The RMS reduction indicates improved behavior for larger errors and outliers.
- Error distribution: Evaluating convergence should include higher norms beyond MAE because they provide more information about outliers and worst-case scenarios.MAE alone can obscure the distribution’s largest errors.
C. Training curves and hyperparameter optimization
Hyperparameter optimization focused on the SOAP cutoff radius, which controls locality and reveals the interaction ranges needed for accurate DFT, CC, and ∆CC-DFT energy learning. Training curves across GDB9 and QM7b show a trade-off between descriptor completeness and extrapolative power for small training sets.
- Training curves: GDB9 DFT training curves used 20,000 FPS-selected structures for training and the remaining 114,000 structures for testing, with shorter cutoffs typically performing better for small ntrain.The error eventually saturates as the cutoff varies.
- Cutoff-radius optimization: Optimization focused on the SOAP cutoff radius rC rather than systematically searching the full parameter space.The cutoff affects test error and indicates the energy scale associated with different degrees of locality.
- Cutoff-radius optimization: For DFT and CC energies, the energy scale decreases from about 3 kcal/mol at rC = 2 ˚A to below 1 kcal/mol at rC = 3 ˚A.Both energy types exhibit a similar trend as the cutoff range increases.
- Cutoff-radius optimization: For ∆CC-DFT energies, rC = 3.5 ˚A is optimal for achieving accuracy below 0.2 kcal/mol because longer-range interactions are crucial.The ∆CC-DFT energy scale is much lower than the absolute DFT and CC energy scales.
- Training curves: GDB9 DFT-energy training curves used about 33k random test structures and the remaining 100k structures in FPS order, with the cutoff increased to 3.5 ˚A for finer-grained energetics.The error was evaluated as MAE versus the number of training inputs.
- Training curves: QM7b and GDB9 training curves show a trade-off between SOAP descriptor completeness and extrapolative power for small training sets.The same qualitative trend appears across different SOAP cutoff lengths and, for GDB9, resembles the behavior observed for CC energies.
D. DFT-on-DFT benchmarks
DFT-on-DFT benchmarks show that SOAP-GAP can predict molecular energetics accurately from DFT geometries, with multi-scale kernels further reducing errors. The model also extrapolates to unseen conformers and improves DFT-to-CC energy estimates.
- DFT-on-DFT benchmarks: With the same training-set size, the additive SOAP-GAP framework achieves a MAE of 0.4 kcal/mol using a 3 Å cutoff.Training-curve behavior reflects a tradeoff between ultimate accuracy and extrapolative power at small training-set sizes.
- DFT-on-DFT benchmarks: A multi-scale SOAP kernel reduces the MAE consistently across training-set sizes, reaching just 0.26 kcal/mol.The kernel combines SOAP information from 2, 3, and 4 Å length scales.
- Dipeptide conformers: Using DFT-optimized geometries, GDB9-trained models predict relative conformer stability for 500 glutamic-acid and 184 aspartic-acid dipeptide minima.The geometries were re-optimized with the same density-functional protocol used for GDB9.
- DFT-to-CC corrections: The GDB9-trained SOAP-GAP model corrects DFT-to-CC atomization-energy discrepancies and provides some correction to relative conformer energetics.The relative-energy correction is notable because the corresponding conformer data were not explicitly included in GDB9.
- Glucose conformers: For glucose, training on 20 FPS-selected structures and validating on 188 others yields energy predictions on par with the most accurate electronic-structure methods.The dataset contains 208 conformers spanning closed and open-chain configurations.