Source-linked AI summary

Electronic Spectra from TDDFT and Machine Learning in Chemical Space

Raghunathan Ramakrishnan, Mia Hartmann, Enrico Tapavicza, O. Anatole von Lilienfeld

arXiv:1504.01966v2physics.chem-ph

TL;DR

TDDFT enables high-throughput spectral prediction but can fail qualitatively for charge-transfer and double-excitation transitions. The paper applies deviation-based machine learning to estimate CC2-level excitation energies for 22 k organic molecules, finding that systematic error analysis can further improve predictions.

  • Problem

    TDDFT offers efficient high-throughput spectral prediction, yet can be qualitatively inaccurate for charge-transfer and double-excitation transitions.

  • Method

    The study trains machine-learning models on deviations between CC2 and TDDFT excitation energies across a chemical universe of 22 k organic molecules.

  • Results

    22 k organic molecules were analyzed, and CC2-level valence excitation energies were estimated at the speed of small-basis-set TDPBE0 through statistical inference of baseline errors.

  • Takeaways & Limitations

    Statistical learning can rectify severe TDDFT prediction flaws, while accounting for systematic chromophore-dependent shifts further improves predictions.

  • Takeaways & Limitations

    Modeling oscillator strengths may fail because the target property requires knowledge of a particular combination of two wavefunctions.

Abstract

from arXiv · show

Due to its favorable computational efficiency time-dependent (TD) density functional theory (DFT) enables the prediction of electronic spectra in a high-throughput manner across chemical space. Its predictions, however, can be quite inaccurate. We resolve this issue with machine learning models trained on deviations of reference second-order approximate coupled-cluster singles and doubles (CC2) spectra from TDDFT counterparts, or even from DFT gap. We applied this approach to low-lying singlet-singlet vertical electronic spectra of over 20 thousand synthetically feasible small organic molecules with up to eight CONF atoms. The prediction errors decay monotonously as a function of training set size. For a training set of 10 thousand molecules, CC2 excitation energies can be reproduced to within $\pm$0.1 eV for the remaining molecules. Analysis of our spectral database via chromophore counting suggests that even higher accuracies can be achieved. Based on the evidence collected, we discuss open challenges associated with data-driven modeling of high-lying spectra, and transition intensities.

I. INTRODUCTION

The paper motivates machine-learning correction of efficient but unreliable TDDFT spectra, using CC2 as a more accurate target and molecular-similarity models of baseline errors.

  • I. INTRODUCTION: TDDFT enables high-throughput spectral prediction but can fail qualitatively for charge-transfer and double-excitation character.These failures reduce its usefulness for screening molecules without inspecting transition orbitals.
  • I. INTRODUCTION: Machine learning offers an alternative for navigating chemical space by inferring quantum-mechanical properties from computed examples.Prior work established accurate interpolation for many ground-state molecular properties.
  • I. INTRODUCTION: The Δ-ML model estimates CC2 excited-state properties as a baseline prediction plus a similarity-weighted correction learned from training molecules.The correction models the baseline method’s deviation from CC2 rather than the target property directly.
  • I. INTRODUCTION: The study models excitation energies and oscillator strengths for the lowest two singlet excited states, using DFT or TDDFT baselines and CC2 as targetline.CC2 is selected as a compromise between accuracy and computational cost.
  • I. INTRODUCTION: Regression coefficients are obtained with a kernel model whose exponential similarity function uses the Manhattan norm between molecular descriptors and regularized least squares.The kernel width and regularization strength control the similarity model and its fit.

B. Cross-validation

Hyperparameters are optimized by five-fold cross-validation, while larger training sets motivate a cheaper single-kernel alternative with stronger assumptions.

  • B. Cross-validation: Five-fold cross-validation partitions the training molecules into five bins, using each bin once for validation and four bins for training.The procedure optimizes σ and λ to minimize validation-bin MAE.
  • B. Cross-validation: The cross-validation procedure requires O(n^3) matrix inversions and takes roughly one CPU day for a fully converged model at the largest training size.Here n is at most 8 k.
  • B. Cross-validation: For larger training sets, the property-independent single-kernel Ansatz avoids prohibitive cross-validation by estimating hyperparameters from training structures alone.It assumes the training data contain no outliers and fixes λ to a property-independent value typically near zero.

C. Choice of molecular descriptor

The study compares Coulomb-matrix and bag-of-bonds molecular representations, finding similar large-data accuracy but a typical small-data advantage for BOB.

  • C. Choice of molecular descriptor: The models use either a row-sorted Coulomb matrix or its more compact bag-of-bonds variant as molecular representations.Sorting the Coulomb matrix by row norms provides invariance to permutations of identical nuclei.
  • C. Choice of molecular descriptor: The Coulomb-matrix off-diagonal elements encode molecular geometry and atomic composition through interatomic distances and atomic numbers.Its diagonal elements approximate the negative potential energy of neutral atoms.
  • C. Choice of molecular descriptor: BOB groups off-diagonal Coulomb-matrix elements by atom-type pairs, enabling pairwise distance comparisons between atom types.This representation is not unique for homometric molecules with identical stoichiometry.
  • C. Choice of molecular descriptor: For large training sets, CM and BOB converge toward similar accuracy for energy-related properties, whereas BOB typically performs better with smaller sets.The dataset contains no homometric molecules, although BOB would assign them zero descriptor difference.

D. Excited states data

The dataset comprises synthetically feasible small organic molecules evaluated with DFT, TDDFT, and RI-CC2 calculations, with several convergence and emission exclusions.

  • D. Excited states data: 134 k GDB-17 molecules with up to 9 CONF atoms were filtered by removing 3,054 molecules for high steric strain.The study then uses a 22 k-molecule chemical universe, as described in the paper context.
  • D. Excited states data: Figure 1 displays joined distributions of first-transition oscillator strengths and energies for the 22 k molecules at the RICC2/def2TZVP level.Representative molecules with large oscillator strengths are shown for selected transition energies.
  • D. Excited states data: The calculations include TDDFT with PBE0/def2SVP, RI-CC2 with def2TZVP, and larger-basis TDDFT using PBE0 and CAM-B3LYP.These calculations provide the baseline and reference spectra used in the study.
  • D. Excited states data: Seven molecules lacked converged RI-CC2 first-excited-state wavefunctions, while seven others showed negative lowest transition energies attributed presumably to orbital relaxation.The calculations were performed with C1 symmetry.

III. RESULTS AND DISCUSSION

The molecular database spans primarily UV-B and UV-C excitation energies, with oscillator strengths generally declining as energy increases. TDDFT and CC2 distributions differ systematically, motivating machine-learning corrections.

  • 22 k molecules span mainly the UV-B and UV-C regions, with few transitions in UV-A.The first-excitation distribution is bimodal, and higher-energy molecules tend to be more saturated.
  • About a dozen molecules have f1 > 0.5, corresponding to light-harvesting efficiencies above 68%.These molecules contain ketoxime or amidine chromophores and show push-pull π-bond conjugation.
  • TDDFT depletes transition-energy densities near 7 eV relative to CC2 and overestimates densities in low- and high-energy regions.The discrepancy is observed for both the first and second singlet transitions.

B. ML models of excitation energies

Delta-ML models correct TDDFT excitation-energy distributions using out-of-sample predictions, and their errors decrease systematically with training-set size. More sophisticated baselines reduce the error offset, while learning rates converge across models.

  • 1 k and 5 k delta-ML models use TDPBE0 baselines to correct both excitation energies on the remaining out-of-sample molecules.The models are trained on randomly selected subsets of the 22 k molecule dataset.
  • Increasing training-set size systematically lowers the out-of-sample MAE for E1 across the evaluated baseline methods.Figure 3 compares errors over training sizes from N = 0 through 10 k.
  • 0.4, 0.3, and 0.13 eV are the 10 k errors for zero, HOMO-LUMO-gap, and TD baselines with def2SVP, respectively.Using def2TZVP improves the PBE0 baseline result to 0.08 eV for 10 k delta-ML.
  • Baseline choice mainly changes the error offset, because the models converge toward similar learning rates as training data increases.The reported learning rate is the slope on the log-log plot of error versus training-set size.

C. Inclusion of systematic shifts

TDDFT errors separate by chromophore type, enabling systematic shift corrections before machine learning. This correction lowers the 10 k E1 MAE to 0.1 eV, while ML also improves severe DFT outliers.

  • −0.31 and +0.19 eV are the signed E1 error centers for saturated σ and unsaturated π chromophores, respectively.The two error distributions are well separated in the 22 k molecule set.
  • 10 k training molecules reduce the E1 MAE from 0.13 eV to 0.1 eV after subtracting chromophore-specific shifts.Saturation can be detected from SMILES strings, allowing the shifts to be applied before delta-ML correction.
  • 2.15 eV is the largest reported TDPBE0 deviation from CC2 among the ten most extreme DFT outliers.These worst DFT outliers are unsaturated molecules.
  • The 5 k ML model improves every listed outlier prediction over DFT, while the 1 k model improves all but one.The exception for the 1 k model is cyclopenta-1-en-4-on.
  • The identities and rankings of the ten most extreme ML outliers are not conserved from the DFT outlier set.A saturated molecule, CF4, appears among the 10 k model outliers with a −1.27 eV deviation.

E. ML models of oscillator strengths

The ∆-ML approach does not systematically improve oscillator-strength predictions with larger training sets, unlike excitation-energy models. Potential explanations include the complexity of oscillator strengths, state-ordering mismatches, and an ill-posed training problem.

  • Oscillator-strength ∆-ML models for f1 and f2 do not become more accurate as training-set size increases.The properties correspond to the S0 → S1 and S0 → S2 transitions, respectively.
  • 0.0101 a.u. is the MAE of TDCAM-B3LYP/def2TZVP oscillator strengths relative to CC2/def2TZVP for the 22 k dataset.
  • The approach may require substantially larger training sets because oscillator strengths depend on a particular combination of two wavefunctions.
  • Different TDDFT and CC2 state orderings can make the baseline and target correspond to different matrix elements, reducing ML-training efficiency.
  • Direct ML models with zero baseline also show insignificant improvement as training-set size increases, so state ordering alone does not explain the failure.

IV. CONCLUSIONS

The study applies ∆-ML to low-lying excitation energies for 22 k organic molecules and finds that statistical learning can correct severe TDDFT prediction flaws. Systematic chromophore-dependent shifts can further improve predictions, while oscillator-strength modeling remains an open challenge.

  • IV. CONCLUSIONS: 22 k organic molecules with up to 8 CONF atoms were used to model low-lying valence spectra at TDDFT and CC2 levels.
  • IV. CONCLUSIONS: Large-basis-set CC2 excitation energies can be estimated at the speed of small-basis-set TDPBE0 by learning deviations from the baseline.
  • IV. CONCLUSIONS: TDPBE0 overestimates the lowest two transition energies for σ-chromophores and underestimates them for π-chromophores.
  • IV. CONCLUSIONS: Accounting for these systematic chromophore-dependent shifts improves ∆-ML predictions and incorporates expert knowledge of error distributions.
  • IV. CONCLUSIONS: Poor oscillator-strength prediction remains a challenge, while the 22 k excited-state database can support benchmarking and chromophore–auxochrome analysis.

APPENDIX: DERIVATION OF EQ. 3 IN MATRIX NOTATION

The appendix derives the kernel-ridge regression linear system by expanding a regularized least-squares objective and setting its derivative with respect to the coefficient vector to zero.

  • APPENDIX: DERIVATION OF EQ. 3 IN MATRIX NOTATION: The derivation starts from a regularized least-squares error measure for estimated training properties pest = Kc and reference values x = pref.
  • APPENDIX: DERIVATION OF EQ. 3 IN MATRIX NOTATION: The residual term is expanded as (x − Kc)^T(x − Kc), with λc^TKc penalizing the fit coefficients.
  • APPENDIX: DERIVATION OF EQ. 3 IN MATRIX NOTATION: The linear system is obtained by setting the derivative of the Lagrangian with respect to the regression-coefficient vector c to zero.
  • APPENDIX: DERIVATION OF EQ. 3 IN MATRIX NOTATION: The derivation uses the symmetry of the kernel matrix, KT = K, together with a matrix-calculus identity.

VI. SUPPLEMENTARY INFORMATION

The supplementary information provides the indices needed to retrieve geometries for the 22 k GDB-8 molecules and collects their TDDFT and CC2 excitation energies.

  • VI. SUPPLEMENTARY INFORMATION: Indices of the 22 k GDB-8 molecules are provided for retrieving their geometries from the 134 k GDB-9 dataset.
  • VI. SUPPLEMENTARY INFORMATION: The supplementary file collects TDDFT and CC2 excitation energies for these molecules.
Loading 1504.01966v2…