Source-linked AI summary
Accurate molecular polarizabilities with coupled-cluster theory and machine learning
David M. Wilkins, Andrea Grisafi, Yang Yang, Ka Un Lao, Robert A. DiStasio, Michele Ceriotti
TL;DR
Accurate molecular polarizabilities are difficult to obtain because the response is highly sensitive to electronic structure, while coupled-cluster calculations are expensive. The paper combines LR-CCSD reference calculations with symmetry-adapted machine learning to predict polarizability tensors efficiently. ALPHA-ML reaches substantially lower error than DFT relative to CCSD and extrapolates to larger molecules, though accuracy worsens for some highly conjugated and sulfur-containing systems.
Problem
Molecular polarizability is important for interactions, spectroscopy, and force fields, but its sensitivity to electronic structure makes accurate prediction difficult and high-level calculations costly.
Method
The study computes LR-CCSD references for QM7b molecules and trains ALPHA-ML using symmetry-adapted λ-SOAP kernels and SA-GPR.
Results
Using 75% of QM7b, the model achieves 2.5% RMSE relative to σCCSD, compared with 18% for DFT and 3.2% when learning DFT polarizabilities.
Takeaways & Limitations
ALPHA-ML provides an inexpensive route to polarizabilities with at least DFT accuracy and often better, while supporting predictions for larger molecules and atom-centered interpretation.
Takeaways & Limitations
Predictions show larger errors for highly conjugated, highly polarizable compounds and sulfur-containing structures that are poorly represented in QM7b.
Abstract
from arXiv · showhide
The molecular polarizability describes the tendency of a molecule to deform or polarize in response to an applied electric field. As such, this quantity governs key intra- and inter-molecular interactions such as induction and dispersion, plays a key role in determining the spectroscopic signatures of molecules, and is an essential ingredient in polarizable force fields and other empirical models for collective interactions. Compared to other ground-state properties, an accurate and reliable prediction of the molecular polarizability is considerably more difficult as this response quantity is quite sensitive to the description of the underlying molecular electronic structure. In this work, we present state-of-the-art quantum mechanical calculations of the static dipole polarizability tensors of 7,211 small organic molecules computed using linear-response coupled-cluster singles and doubles theory (LR-CCSD). Using a symmetry-adapted machine-learning based approach, we demonstrate that it is possible to predict the molecular polarizability with LR-CCSD accuracy at a negligible computational cost. The employed model is quite robust and transferable, yielding molecular polarizabilities for a diverse set of 52 larger molecules (which includes challenging conjugated systems, carbohydrates, small drugs, amino acids, nucleobases, and hydrocarbon isomers) at an accuracy that exceeds that of hybrid density functional theory (DFT). The atom-centered decomposition implicit in our machine-learning approach offers some insight into the shortcomings of DFT in the prediction of this fundamental quantity of interest.
I. INTRODUCTION
Molecular polarizability is a fundamental response property that is difficult to predict accurately because it is sensitive to electronic-structure details, while LR-CCSD accuracy is computationally expensive. The paper benchmarks coupled-cluster polarizabilities and develops ALPHA-ML to predict them efficiently for small and larger molecules.
- Molecular polarizability governs induction and dispersion interactions and contributes to spectroscopic signatures and polarizable force fields.
- Accurate polarizabilities require simultaneous treatment of electron-correlation effects and basis-set incompleteness because polarizability is a sensitive response property.
- LR-CCSD with sufficiently diffuse basis sets provides more reliable polarizabilities than DFT but scales with the sixth power of system size, limiting applications to moderate molecules.
- The study uses 7,211 QM7b molecules to benchmark coupled-cluster polarizabilities, assess hybrid DFT, and train the symmetry-adapted ALPHA-ML model.
- B3LYP errors relative to LR-CCSD are 0.302 a.u. MAE and 0.404 a.u. RMSE, while SCAN0 reduces systematic error but retains large statistical errors.
Improved Symmetry-Adapted Gaussian Process
The improved symmetry-adapted Gaussian-process approach accelerates tensorial polarizability learning while preserving the required covariance properties. Its main efficiency gains come from sparsifying λ-SOAP representations and adding controlled nonlinear kernels.
- Sparsifying λ-SOAP representations with a few hundred spherical-harmonic components preserves essentially the full-kernel result at much lower computational cost.
- The model introduces nonlinearity by multiplying λ > 0 kernels by a scalar λ = 0 kernel raised to the power ζ − 1.
- This kernel construction increases the represented interatomic body order without compromising the tensorial covariance required by the model.
Learning on the QM7b Database
ALPHA-ML learns CCSD molecular polarizabilities from QM7b with substantially lower error than DFT, while also supporting correction-based learning between theory levels.
- Model training: ALPHA-ML is trained on QM7b reference calculations to predict molecular polarizabilities using a transferable machine-learning model.The learning assessment uses up to 5,400 training structures and 1,811 held-out molecules.
- CCSD learning: 2.5% RMSE relative to σCCSD is achieved when 75% of QM7b is used to predict CCSD polarizability.The intrinsic variability is reported as σCCSD = 2.216 a.u. per atom.
- Comparison with DFT: 18% of σCCSD is the DFT error relative to CCSD, compared with 2.5% for machine-learning prediction of CCSD polarizability.The DFT polarizability itself can be learned with a 3.2% error relative to σCCSD.
- Correction learning: Learning the CCSD–DFT correction reduces the error by about a factor of 2 relative to directly learning αCCSD.This correction strategy requires a baseline DFT calculation.
- Reference-method robustness: ALPHA-ML shows similar accuracy when trained on SCAN0 and B3LYP reference calculations.The paper therefore focuses its main discussion on B3LYP and LR-CCSD results.
Extrapolation to Larger Molecules
ALPHA-ML is evaluated on 52 larger molecules beyond QM7b and retains better-than-DFT accuracy, although performance worsens for highly delocalized and poorly represented chemical systems.
- Showcase dataset: The showcase dataset contains 52 larger molecules, including amino acids, nucleobases, drugs, carbohydrates, alkanes, and alkenes.The model is trained on the full QM7b database before this extrapolative test.
- Accuracy comparison: ALPHA-ML predicts CCSD polarizabilities more accurately than DFT for the showcase molecules.Table I reports RMSEs for CCSD/DFT, CCSD/ML, DFT/ML, and learning the CCSD–DFT difference, including scalar and tensorial components.
- Baseline correction: 20–30% further error reduction is obtained when DFT is used as a baseline for ALPHA-ML prediction of CCSD polarizability.The model learns both λ = 0 and λ = 2 components with similar efficiency.
- Extrapolation boundary: Showcase performance is considerably worse than validation performance on the QM7b database, despite remaining better than DFT.This comparison marks the principal extrapolation boundary reported for the model.
- Error patterns: Most showcase molecules have small errors, but large errors occur mainly for highly polarizable, strongly conjugated compounds and sulfur-containing structures.Long-chain alkenes and purine nucleobases are identified as especially challenging, while sulfur-containing structures are poorly represented in QM7b.
- Error patterns: The tensorial α(2) predictions are generally slightly less accurate than DFT, except for highly polarizable alkenes where ALPHA-ML performs dramatically better.The difficult cases are associated with electronic delocalization and require larger cutoffs and more complex reference molecules.
Atom-Centered Environmental Polarizabilities
ALPHA-ML decomposes molecular polarizability into atom-centered contributions that can be interpreted through local chemical environments, while also exposing systematic DFT–CCSD differences.
- Atom-centered decomposition: ALPHA-ML represents the molecular polarizability as a sum of atom-centered local terms produced by the SA-GPR model.These contributions arise as a byproduct of the model and require no additional calculations.
- Interpretation and caveats: The decomposition is inductive rather than a literal atoms-in-molecules partition, so individual atomic contributions can have negative eigenvalues.Several hydrogen environments show negative eigenvalues, indicating contributions that reduce the dielectric response.
- Chemical interpretation: Large, anisotropic contributions are assigned to unsaturated carbon and aromatic systems, consistent with electron delocalization in conjugated regions.Examples include octatetraene, the indole ring of tryptophan, and adenine.
- Chemical interpretation: The atom-centered decomposition helps explain transferability to larger molecules by assigning functional-group contributions from relatively local structural information.The paper connects this interpretability to ALPHA-ML's better-than-DFT performance on larger showcase molecules.
- DFT error analysis: Difference learning attributes positive CCSD–DFT corrections to carbon atoms and negative corrections to oxygen atoms.This pattern suggests that DFT overestimates carbon-centered polarizability and underestimates polarizability in oxygen-containing moieties.
DISCUSSION
CCSD improves polarizability accuracy over DFT but can be prohibitively expensive, motivating ALPHA-ML as a lower-cost alternative that retains at least DFT-level accuracy and often performs better.
- DISCUSSION: CCSD polarizability calculations improve accuracy over DFT but incur a steep computational-cost increase.For very large showcase molecules, LR-CCSD calculations required up to hundreds of hours and thousands of gigabytes of RAM.
- DISCUSSION: ALPHA-ML combines SA-GPR, λ-SOAP kernels, and small-molecule CCSD references to avoid repeated expensive calculations.The framework predicts polarizabilities with at least DFT accuracy, often substantially better, at a fraction of the computational cost.
- DISCUSSION: The model extrapolates from small training molecules to larger and more complex compounds, with prediction quality that can improve as the training set expands.The authors identify extension to larger compounds, oligomers, and condensed-phase systems as future work.
- DISCUSSION: Future extensions target polarizable force fields and computational predictions of Raman scattering and sum-frequency generation spectroscopy.These applications are framed as potential outcomes of extending the model to larger systems.
Polarizabilities in the QM7b Database
The study evaluates polarizabilities using finite-field DFT and LR-CCSD calculations on QM7b and larger showcase molecules, with finite-field CCSD used for eight especially demanding cases.
- Computational setup: The 52 showcase geometries were optimized with PBE in FHI-AIMS, while QM7b learning geometries were taken from the database.The calculations used Psi4 for B3LYP and LR-CCSD, and Q-Chem for SCAN0.
- DFT polarizabilities: Finite-field DFT obtains polarizabilities as numerical derivatives of dipole moments with respect to applied electric fields.The field perturbation was applied along x, y, and z using δE = 1.8897261250 × 10^-5 a.u.
- CCSD polarizabilities: Most CCSD polarizabilities were calculated directly with LR-CCSD singles and doubles.The calculations used a frozen-core approximation and direct SCF treatment.
- CCSD polarizabilities: Eight large showcase molecules required an orbital-non-relaxed finite-field CCSD calculation because of RAM and disk-memory demands.For these molecules, polarizability was obtained from the second derivative of perturbed CCSD energies.
- Computational setup: All calculations used the d-aug-cc-pVDZ Dunning-type basis set with method-specific convergence settings.The basis was obtained from the EMSL Basis Set Library.
Error Assessment
The study evaluates polarizability predictions with rotation-invariant Frobenius-norm errors and reports per-atom statistical measures to compare molecules of different sizes.
- Error Assessment: The Frobenius norm measures errors while including scalar and anisotropic polarizability components and remaining invariant to rigid rotations.The metric compares two sets of polarizability tensors.
- Error Assessment: Mean signed error and related statistical measures are defined over N molecular structures containing n_i atoms.The framework compares predicted and reference polarizabilities across the dataset.
- Error Assessment: Errors are defined on a per-atom basis to simplify comparisons between datasets containing molecules of different sizes.This normalization accounts for variation in molecular atom counts.
Symmetry-Adapted Gaussian Process Regression
The SA-GPR model decomposes polarizability tensors into symmetry-adapted components and learns them from atom-centered λ-SOAP descriptors and kernels. The workflow uses sparse representations and separate kernel-ridge regressions while preserving the required symmetry properties.
- Tensor decomposition: Each polarizability tensor is decomposed into irreducible real spherical components, comprising one scalar component and a 5-vector.The Cartesian-to-spherical transformation is unitary, so the Frobenius norm is preserved.
- Environment representation: λ-SOAP descriptors encode interatomic correlations within a prescribed cutoff around each atom-centered environment.These descriptors are used to learn tensor components of order λ.
- Kernel construction: The linear SOAP kernel captures atomic correlations up to 3-body terms, while normalization and integer powers introduce many-body correlations.The λ-SOAP kernel uses selected spherical-harmonic descriptor components.
- Sparsification: A farthest-point sampling procedure selects a subset of spherical-harmonic components to form a potentially sparse representative basis.The representative environments are used as the basis for the kernel model.
- Regression and symmetry: Separate kernel-ridge regression models determine weights for each polarizability component.The λ-SOAP kernels must retain their linear nature to enforce the correct symmetry properties.