Source-linked AI summary

Accurate and scalable exchange-correlation with deep learning

Giulia Luise, Chin-Wei Huang, Thijs Vogels, Derk P. Kooi, Sebastian Ehlert, Stephanie Lanius, Klaas J. H. Giesbertz, Amir Karton, Deniz Gunceler, Stefano Battaglia, Gregor N. C. Simm, P. Bernát Szabó, Megan Stanley, Wessel P. Bruinsma, Lin Huang, Xinran Wei, José Garrido Torres, Abylay Katbashev, Rodrigo Chavez Zavaleta, Bálint Máté, Sékou-Oumar Kaba, Roberto Sordillo, Yingrong Chen, David B. Williams-Young, Christopher M. Bishop, Jan Hermann, Rianne van den Berg, Paola Gori-Giorgi

arXiv:2506.14665v6physics.chem-phcs.AIcs.CEcs.LGphysics.comp-ph

TL;DR

DFT predictions remain limited by exchange-correlation approximations that do not reliably achieve experimental accuracy across chemical systems and properties. Skala uses large-scale wavefunction data and a non-local deep-learning architecture with semi-local inputs, outperforming leading hybrid functionals across broad main-group chemistry at semi-local DFT cost. The results show that deep learning can break the traditional accuracy–cost trade-off and improve systematically as training data expand.

  • Problem

    Current exchange-correlation approximations do not reliably achieve chemical accuracy across diverse chemical systems and properties, while machine-learning functionals have not shifted the accuracy–cost trade-off.

  • Method

    Skala combines approximately 400,000 high-accuracy energy differences with a neural architecture that learns non-local electronic representations from simple semi-local inputs.

  • Results

    Skala achieves higher accuracy than leading hybrid functionals across broad main-group chemistry while retaining the computational cost and scaling of semi-local DFT.

  • Takeaways & Limitations

    Deep learning enables an exchange-correlation functional framework whose accuracy improves systematically as high-accuracy training data expand without increasing inference complexity.

  • Takeaways & Limitations

    Extending Skala to transition-metal chemistry requires substantially expanding the scale and diversity of high-accuracy training data, with multi-reference chemistry remaining a challenge.

Abstract

from arXiv · show

Density Functional Theory (DFT) underpins much of modern computational chemistry and materials science. Yet, the reliability of DFT-derived predictions of experimentally measurable properties remains fundamentally limited by the need to approximate the unknown exchange-correlation (XC) functional. The traditional paradigm for improving accuracy has relied on increasingly elaborate hand-crafted functional forms. This approach has led to a longstanding trade-off between computational efficiency and accuracy, which remains insufficient for reliable predictive modelling of laboratory experiments. Here we introduce Skala, a deep learning-based XC functional that surpasses state-of-the-art hybrid functionals in accuracy across the main-group chemistry benchmark set GMTKN55 with an error of 2.8 kcal/mol, while retaining the lower computational cost characteristic of semi-local DFT. This demonstrated departure from the historical trade-off between accuracy and efficiency is enabled by learning non-local representations of electronic structure directly from data, bypassing the need for increasingly costly hand-engineered features. Leveraging an unprecedented volume of high-accuracy reference data from wavefunction-based methods, we establish that modern deep learning enables systematically improvable neural exchange-correlation models as training datasets expand, positioning first-principles simulations to become progressively more predictive.

1 Introduction

DFT is widely used for predictive chemistry and materials modelling, but existing exchange-correlation approximations do not reliably achieve chemical accuracy across chemical systems. Skala addresses this longstanding accuracy–cost trade-off with deep learning, large-scale high-accuracy data, and systematic improvement as training data expand.

  • Motivation: Chemical accuracy generally requires errors below ∼1 kcal/mol, yet current exchange-correlation approximations remain unreliable across diverse chemical systems and properties.This limitation can cause screening pipelines to send too many candidates for experimental verification.
  • Accuracy–cost trade-off: Traditional DFT advances handcraft increasingly elaborate functionals along Jacob’s ladder, trading semi-local efficiency for more expensive wavefunction-based information.Lower-rung methods retain asymptotic O(N^3) scaling but use only semi-local electronic information.
  • Machine-learning limitations: Machine-learning functionals have not shifted this trade-off because accurate wavefunction-based training data are scarce and difficult to generate at scale.The low-data regime has limited most efforts and prevented widespread adoption of ML-based functionals.
  • Skala: Skala uses an efficient wavefunction-based protocol to assemble approximately 400,000 high-accuracy energy differences spanning diverse chemistry, supplemented by comparable public datasets.The training strategy directly addresses the scarcity of chemically accurate reference data.
  • Skala: As increasingly diverse high-accuracy data are added, Skala’s accuracy improves systematically without increasing inference-time computational complexity.This scaling behaviour motivates Skala as a continuously evolving exchange-correlation modelling framework rather than a static functional.

2 Deep learning the XC functional: basic challenges and solutions

Skala learns the XC functional with a neural network designed for deployment inside self-consistent-field calculations. Its training and architecture address SCF instability, exact constraints, scalable non-locality, and the need for large high-accuracy datasets.

  • Neural XC learning: Skala parameterizes the XC functional Eθ_xc[ρ] with a neural network whose predictions combine with known energy terms to estimate total energies.The learned functional is evaluated within an SCF loop, which minimizes total energy with respect to the density.
  • Training protocol: A two-stage protocol pre-trains Skala on fixed input densities, then fine-tunes it using self-consistent densities generated during training.Pre-training avoids the computational cost and convergence failures of running SCF calculations for every example from a poor initial functional.
  • Architecture and constraints: Skala incorporates energetically relevant exact constraints and learns non-local interactions between grid points through atom-centred coarse points while preserving scalable computational complexity.The architecture begins from meta-GGA semi-local features and exchanges information between each grid point and its atom’s coarse point.
  • Accuracy–cost trade-off: Learned non-locality removes the need for computationally expensive higher Jacob’s ladder ingredients, enabling high main-group accuracy without increasing computational cost and allowing systematic improvement with more data.The stated targets are main-group thermochemistry and kinetics, rather than dispersion interactions.
  • Training data: 395,000 energy differences in MSR-ACC are labeled with W1-F12 and W1w protocols achieving CCSD(T)/CBS accuracy, including roughly 120,000 total atomization energies.The datasets cover broad chemical space, including affinities, ionization potentials, reaction pathways, and non-covalent clusters.

3 Accuracy of Skala

Skala achieves leading accuracy across main-group chemistry while retaining semi-local computational cost, including chemical accuracy on most W4-17 atomization energies. Its performance improves substantially through learned non-locality and broader training data, while the full dataset also yields physically consistent kinetic-correlation behavior.

  • GMTKN55 accuracy: Skala achieves the lowest GMTKN55 WTMAD-2 error, surpassing all compared hybrid functionals at semi-local cost.GMTKN55 is the standard main-group chemistry benchmark used for comparison across widely used functionals.
  • GMTKN55 accuracy: 32 of 55 GMTKN55 subsets are best-performing for Skala, exceeding all other functionals combined and covering nearly all thermochemistry and kinetics benchmarks.The advantage extends from small to large systems and challenges the assumption that this accuracy requires hybrid or double-hybrid functionals.
  • W4-17 atomization energies: 0.92 kcal/mol MAE is achieved on the single-reference W4-17 subset, while overall W4-17 MAE is 1.23 kcal/mol across all 200 reactions.The single-reference subset contains 183 reactions; the multi-reference cases lack corresponding training data.
  • Ablations: About 50% lower average WTMAD-2 error across 5 seeds results from adding Skala’s non-local branch compared with purely local models.The models share the same architecture apart from non-locality and are evaluated on Diet GMTKN55.
  • Ablations: Systematic accuracy improvement on Diet GMTKN55 follows the addition of reactions, barrier heights, non-covalent interactions, and conformers to the training data.Set A reaches GGA-level accuracy, set B provides a modest improvement, and set C adds further improvement.
  • Physical constraints: Training on the full dataset consistently satisfies the positivity constraint for Tc, unlike training only on MSR-ACC/TAE25.The resulting Tc values are not far from reverse-engineered CCSD-density values, although the GKS framework permits differences from pure Kohn-Sham Tc.

4 Beyond energies: Densities and equilibrium geometries

Skala’s self-consistent-density fine-tuning is evaluated through dipole moments to test density quality, while equilibrium-geometry benchmarks assess its structural accuracy. It reaches at least hybrid-functional-level geometry accuracy, improving substantially over training limited to equilibrium structures and reaction paths.

  • Densities: Fine-tuning Skala on self-consistent densities closes the gap between accuracy learned on B3LYP densities and accuracy obtained on Skala’s self-consistent densities.The procedure pretrains on wavefunction energy differences evaluated on ρB3LYP, then fine-tunes for a small number of steps using on-the-fly SCF densities.
  • Densities: After ∼2000 fine-tuning steps, all systems converge and the reported dipole and energy errors reflect the full dataset.During the initial ≲2000 steps, SCF convergence is erratic and errors are reported only for converged molecules.
  • Equilibrium geometries: Training only on equilibrium structures and reaction paths yields, at best, GGA-level accuracy for predicted equilibrium geometries.This motivates augmenting training data to target specific regions of chemical space and properties.
  • Equilibrium geometries: Skala geometries are benchmarked against LMGB35, HMGB11, CCse21, and CCSD(T)/CBS W4-11-GEOM datasets, alongside functionals across Jacob’s ladder and GFN2-xTB.The benchmarks comprise light and heavy main-group bond lengths plus bond lengths and angles for small molecules.
  • Equilibrium geometries: Skala achieves equilibrium-geometry accuracy on par with, or exceeding, that of the best hybrid functionals.The comparison also includes the semi-empirical GFN2-xTB method.

5 Computational Efficiency and Integration into Production Codes

Skala is designed to retain meta-GGA semi-local DFT’s asymptotic scaling, but its practical efficiency requires empirical validation because prefactors and bottlenecks vary across methods and system sizes. GPU measurements support scalable semilocal-like cost, while grid-level non-locality creates an integration barrier for production-code libraries.

  • Computational efficiency: Hardware-specific optimizations and algorithmic advances continue to lower computational cost in practice.This reinforces why asymptotic scaling alone does not determine realized efficiency.
  • Computational efficiency: Skala has the same asymptotic scaling as meta-GGA semi-local DFT, but its prefactor and actual cost must be empirically verified as system size increases.Practical costs can differ substantially because prefactors, dominant terms, and memory bottlenecks vary with method and molecular size.
  • Computational efficiency: 30%: On GPU, Skala’s measured XC-integration cost with sn-LinK falls within 30% of the truncated comparison reported for the tested functionals.The measurements isolate XC evaluation from other components of a Kohn–Sham calculation and include exact exchange for B3LYP and M06-2X.
  • Integration into production codes: Skala’s grid-level non-locality is incompatible with production-code libraries such as libXC, which process grid points independently.This incompatibility is identified as a barrier to widespread adoption despite the architecture’s confirmed scalable semilocal-like cost.

6 Conclusions

Skala is presented as a deep learning-based exchange-correlation functional that learns non-local quantum mechanical effects from semi-local inputs while retaining semi-local DFT’s favourable scaling. The authors argue that expanded training data and the Kohn–Sham framework support systematic improvements and generalization beyond the current main-group focus.

  • Conclusions: Skala is presented as a significant step toward a general-purpose density functional that is chemically accurate and computationally efficient.The conclusion frames this goal as part of the longstanding effort to combine predictive accuracy with computational efficiency.
  • Conclusions: Skala learns non-local quantum mechanical effects from simple semi-local inputs without sacrificing semi-local DFT’s favourable scaling.Its design combines a large-scale training set, an efficiently scaling training protocol, and a non-local architecture with low inference cost.
  • Conclusions: The authors expect Skala’s accuracy and generality to improve systematically as the training dataset expands across broader chemical phenomena.They have begun extending the training set beyond main-group chemistry by incorporating minimal atomic information for 3d and 4d transition metals.
  • Conclusions: Learning an XC functional from high-accuracy wavefunction data allows the Kohn–Sham framework to capture dominant energy contributions relevant to unseen elements and larger systems.The XC functional is described as a smaller correction term, while embedded physical constraints help Skala remain robust.

7 Methods

Skala combines a deep neural network with semi-local meta-GGA density features to model non-local electronic interactions on DFT integration grids. It is trained on approximately 400k high-accuracy reaction energies spanning molecular, non-covalent, kinetic, and transition-metal data.

  • Model architecture: Skala takes semi-local, density-dependent meta-GGA features with O(N^3) scaling as inputs to a deep neural network that models non-local interactions across the integration grid.The features are represented on the large, irregular numerical integration grid used in DFT.
  • Model architecture: Seven log-transformed semi-local inputs are processed by a local MLP for each spin ordering and averaged into a spin-symmetrized hidden representation.This representation is then passed to the non-local interaction model.
  • Model architecture: Non-local information is aggregated at coarse points from fine-grid features projected onto radial basis functions and spherical harmonics.The coarse-point construction is analogous to accumulating multipole moments.
  • Model architecture: 2 is the upper range of the scaled sigmoid enhancement factor, which bounds the output and enforces the Lieb–Oxford lower bound.The final non-local features are processed by a local MLP and projected to one scalar per grid point.
  • Training data: ∼400k reaction energies form the training data, computed at CCSD(T)/CBS level or higher across diverse molecular, conformational, non-covalent, kinetic, and transition-metal subsets.The largest subset contains ∼120k diverse total atomization energies for molecules with up to nine non-hydrogen atoms.

Method References

The method references span equivariant and atomistic machine-learning approaches, foundational density-functional theory, and benchmark datasets for thermochemistry, noncovalent interactions, and reaction energetics. They also document computational frameworks, basis sets, dispersion corrections, and applications to molecular and transition-metal systems.

  • Machine-learning methods: The references include equivariant non-local electron-density functionals, MACE message-passing neural networks, and the Atomic Cluster Expansion.These works represent machine-learning approaches for non-local electronic structure and atomistic modeling.
  • DFT foundations: Foundational references cover exact exchange in density-functional thermochemistry, lower bounds for Coulomb energies, fractional-electron problems, and DFT force fields.These works provide theoretical and methodological context for exchange-correlation modeling and molecular properties.
  • Benchmark datasets: The cited benchmark resources cover GMTKN55, W4-17, W4-11, Accurate Chemistry Collection, GSCDB137, and datasets for gold-standard interaction energies.The references include broad main-group thermochemistry, atomization energies, energy differences, and intermolecular interactions.
  • Benchmark datasets: Noncovalent-interaction references address hydrogen bonding, repulsive contacts, σ-hole interactions, and London dispersion across chemical spaces.The cited Non-Covalent Interactions Atlas datasets comprise five benchmark resources spanning these interaction classes.
  • Computational implementation: The methods infrastructure includes PySCF, Karlsruhe basis sets, minimally augmented basis sets, D3 dispersion damping, and robust atomistic modeling protocols.The references also address applications to materials, organometallic, biochemical, and transition-metal systems.

Supplementary information: … A.2 In practice

Skala models the unknown exchange-correlation functional with a deep neural-network enhancement factor applied to the LDA exchange energy density. In practice, it discretizes semi-local density features on an integration grid and learns non-local dependencies across grid points from data, including the B3LYP D3 correction during training.

  • Supplementary information:: The supplementary information is titled “Accurate and scalable exchange-correlation with deep learning.”
  • A Modeling the exchange-correlation functional: The modeling section defines the input features and neural-network architecture used to represent the exchange-correlation functional.
  • A.1 In theory: DFT expresses ground-state energy through electron density, but its exchange-correlation component Exc[ρ] is unknown and requires accurate approximation.
  • A.1 In theory: Skala parameterizes the enhancement factor fθ with a deep neural network multiplying the LDA exchange energy density, with fθ = 1 recovering LDA exchange.
  • A.2 In practice: In practice, density features are evaluated at G points on a classical integration grid with associated weights to numerically approximate spatial integrals.
  • A.2 In practice: The neural network receives seven semi-local features per grid point, including Kohn-Sham kinetic-energy and two-spin-channel information, as a tensor in R^G×7.
  • A.2 In practice: Non-locality can arise from hand-designed cross-point features or from learning longer-range dependencies across grid points; Skala adopts the latter approach.
  • A.2 In practice: Skala infers non-locality through data and model mixing rather than hand-designed features, while training on B3LYP densities includes the D3 correction and targets shorter-range non-local effects.

A.3 Neural network architecture … B.1 Training objective

Skala combines a three-part atom-centered neural architecture with coarse-point message passing to model non-local, higher-body-order exchange-correlation effects, while training minimizes reaction-energy errors against reference data. Its design yields atom-dependent energy densities and builds on ACE-inspired systematic expressivity.

  • A.3 Neural network architecture: The architecture processes seven semi-local features per atom-centered integration grid and comprises an input extractor, L non-local layers, and an output model.The model uses Dhid = 256, Dnonl = 16, and L = 3 in all experiments.
  • A.3 Neural network architecture: The input representation is evaluated with both spin-channel orderings and averaged to produce spin-order-invariant features.The extractor uses Swish activations and fully connected layers.
  • A.4 Non-local interaction through coarse points: Non-local layers communicate through atom-centered coarse points, aggregating each atom’s grid features, mixing them, and sending processed features back to that atom’s grid points.This avoids all-to-all communication among the large number of integration-grid points.
  • A.4 Non-local interaction through coarse points: Symmetric contraction with ν = 3 extends the layer beyond two-body interactions to up to four-body interactions, while shared weights reflect dependence on electron density rather than nuclear charges.The post-upsampling factor exp(−ρi) suppresses non-local effects near high-density nuclear regions.
  • A.5 Atomic partition of the energy density: The atom-wise grid partition produces atom-dependent energy densities whose contributions are combined using merged integration weights and augmented across partition schemes during training.Each grid point communicates only with its associated atom’s coarse point.
  • A.6 Intuition on the structure of the non-local layer: A single coarse-point downsampling–upsampling pass can approximate two-body interactions, while symmetric tensor products systematically extend representational capacity to higher-body-order interactions.The construction is motivated by rotationally invariant basis expansions around atomic centers.
  • A.7 Related work: The theoretical foundation is based on atomic cluster expansion, with expressivity increased through larger basis-set parameters and higher interaction correlation order; coarse points provide computational and modeling benefits.The non-local layer also shares similarities with low-rank kernel neural operators.
  • B.1 Training objective: Training minimizes a weighted mean squared error on reaction energies, formed as stoichiometric linear combinations of molecular total energies using precomputed B3LYP-based non-XC energies containing D3 dispersion.Reference reaction energies are sampled using the hierarchical procedure described in Sec. B.2.

B.2 Dataset sampling and model selection … C Training data: computational details

The paper balances heterogeneous training data through adaptive dataset sampling, trains the model with specialized initialization and optimization procedures, and applies self-consistent fine-tuning using model-generated densities. It also specifies computational protocols for producing the training data.

  • B.2 Dataset sampling and model selection: Heterogeneous datasets are sampled in two levels: a dataset is selected first, followed by uniform reaction sampling within that dataset.This prevents large datasets from dominating training and underrepresenting smaller chemically important datasets.
  • B.2 Dataset sampling and model selection: Initial dataset probabilities combine category weights with dataset sizes to achieve prescribed sampling proportions across nine chemical categories.The categories include atomization energies, structures, barrier heights, thermochemistry, non-covalent interactions, distortion data, atomic sets, and transition metals.
  • B.2 Dataset sampling and model selection: The relative excess loss measures progress across holdout sets by comparing each dataset’s MAE with baseline minimum and median performance.Ri = 0 indicates performance at least as accurate as the best baseline, while Ri = 1 corresponds to the median baseline level.
  • B.2 Dataset sampling and model selection: Adaptive sampling increases probabilities for datasets with Ri > 0 and downweights datasets where Ri = 0, using η = 0.025 every 25,000 steps.The update redirects computation toward categories that still require improvement.
  • B.3 Parameter initialization: Local network weights use Xavier initialization with zero biases, while non-local layers use a modified Xavier scheme accounting for shared weights among same-order tensor features.The modified scheme is designed specifically for the non-local layer’s tensor-feature structure.
  • B.4 Optimization: Hidden weight matrices use Muon, whereas biases and the final output layer use Adam; Adam uses betas (0.9, 0.95) and ϵ = 10^-10.Muon applies momentum SGD followed by projection of each update to the nearest orthogonal matrix.
  • B.5 Self-consistent fine-tuning: B3LYP-feature functionals can perform worse self-consistently than on B3LYP densities, motivating fine-tuning on self-consistent densities generated by the model itself.Fine-tuning reinitializes optimizer state and uses a constant global learning rate of 10^-5 for both Muon and Adam.
  • C Training data: computational details: Self-consistent densities are produced with PySCF and DIIS over the latest 8 updates, using def2-QZVP densities and convergence thresholds of 10^-8, 5 · 10^-6 Eh, and 1 · 10^-3 Eh.The SCF procedure starts from a B3LYP guess and terminates after 40 steps if convergence is not achieved; the lowest-energy result is retained.

C.1 MSR-ACC/TAE dataset … D Evaluation protocols

The paper constructs MSR-ACC datasets from systematically generated and transformed molecular structures labeled with high-accuracy wavefunction-based protocols, while specifying grid, partitioning, density, and evaluation procedures for training and testing.

  • C.1 MSR-ACC/TAE dataset: MSR-ACC/TAE enumerates plausible covalent graphs for closed-shell, charge-neutral molecules with up to four non-hydrogen atoms through argon, then relaxes their 3D structures.The protocol also samples graphs with up to five non-hydrogen atoms.
  • C.1 MSR-ACC/TAE dataset: ∼41k structures extend MSR-ACC/TAE using GDB9-derived element substitutions, geometry relaxation with GFN2-xTB and B3LYP-D3(BJ)/def2-TZVPP, and W1w-labeled atomization reactions.The substitutions replace second-row elements with third-row elements and hydrogen with Li, Na, F, or Cl.
  • C.2 MSR-ACC/IP, /EA, /PA, /Conf, /NCI, /Water, /Distortion and /Reactions datasets: MSR-ACC’s remaining datasets use reactions labeled with a slightly modified W1w protocol, while ionization, electron-affinity, proton-affinity, conformational, and noncovalent reactions derive from transformed or sampled molecules.The source describes electron removal, electron addition, proton addition, CREST-generated conformational changes up to 10 kcal/mol, and fragment-based noncovalent systems.
  • C.2 MSR-ACC/IP, /EA, /PA, /Conf, /NCI, /Water, /Distortion and /Reactions datasets: W1w extrapolates the Hartree–Fock energy to the complete-basis-set limit from jul-cc-pV(T+d)Z and jul-cc-pV(Q+d)Z using E(L) = E∞ + A/L^α with α = 5.The valence CCSD correlation energy is extrapolated from the same basis sets.
  • C.3 Atomic datasets: Atomic total energies, ionization potentials up to triple ionization, and electron affinities for atoms up to argon were calculated at CCSD(T)/CBS using aug-cc-pCVQZ and aug-cc-pCV5Z extrapolations.Lithium and beryllium are excluded from these atomic datasets.
  • C.4 Density features on the grid: Four distinct integration grids augment the density during training to regularize the model against numerical grid variations.The grids use PySCF level 1 and four radial schemes: Treutler, Mura–Knowles, Gauss-Chebyshev, and Delley.
  • C.5 Augmentation of the space-partitioning schemes during training: Randomly sampling among three space-partitioning schemes during training reduces overfitting to any single atomic-contribution decomposition of the exchange-correlation energy.The molecular exchange-correlation energy should be invariant to the partitioning choice.
  • D Evaluation protocols: Evaluation covers four benchmark categories: reaction energies, dipole moments, geometry optimization, and computational cost.Training evaluates the functional at fixed B3LYP densities using def2-QZVP, or ma-def2-QZVP for reactions containing anions, and computes relative reaction energies from B3LYP total energies without the exchange-correlation energy.

D.1 Evaluation sets … E.3 Convergence with respect to the grid size

The paper evaluates Skala across diverse reaction-energy, dipole-moment, geometry, and computational-cost benchmarks against established semi-local and hybrid functionals. It also specifies convergence safeguards, geometry-gradient implementation, spin-symmetry-breaking tests, and grid-convergence analyses.

  • D.1 Evaluation sets: GMTKN55, W4-17, Wiggle150, and additional holdout datasets assess Skala across reaction energies and diverse chemical systems.GMTKN55 covers basic properties, thermochemistry, kinetics, intermolecular interactions, and conformational energies.
  • D.1 Evaluation sets: Dipole moments, experimental geometries, and a 38-molecule cost set extend evaluation beyond reaction energies.Geometry sets include LMGB35, HMGB11, CCse21, and W4-11-geom; the cost set combines Grimme, S30L, HS13L, and NCI16L systems.
  • D.2 Baselines and dispersion correction settings: Skala is compared with revPBE, B97M-V, r2SCAN, B3LYP, M06-2X, ωB97X-V, and ωB97M-V using specified D3 or VV10 dispersion corrections.revPBE, r2SCAN, B3LYP use D3(BJ), M06-2X uses D3(0), and Skala uses D3(BJ) with B3LYP parameters.
  • D.3 Basis sets: Reaction, dipole, and geometry evaluations use def2-QZVP or ma-def2-QZVP, aug-pc-4, and def2-TZVP basis sets, respectively, with density fitting.The def2-universal-jkfit auxiliary basis is used for all evaluations.
  • D.5 SCF with orbital gradient descent fallback: A gradient-descent fallback with line search improves SCF robustness when DIIS does not reduce energy, using orbital-space updates and adaptive step sizes.Calculations using gradient descent are performed in double precision, with downgraded-basis convergence followed by projection to a larger basis used in some attempts.
  • D.6 Geometry optimization: implementation: Skala geometry optimization includes grid-position dependence and coarse-point contributions in nuclear forces, then uses geomeTRIC with tightened convergence tolerances.The initial geometry uses the SCF retry protocol, while subsequent steps retain the method that successfully converged the initial electronic structure.

E.4 Emergence of exact constraints with more training data

With sufficient data, Skala can learn exact exchange-correlation constraints as emergent behavior without restricting neural-network expressivity. Despite generalized Kohn-Sham complications, it learns a positive correlation kinetic energy close to the pure-Kohn-Sham value across random seeds.

  • Emergent exact constraints: Sufficient training data may enable a neural-network functional to satisfy exact constraints emergently, avoiding restrictions that could limit model expressivity.Traditional approximate functionals impose constraints by restricting their functional form, whereas neural networks may learn them from data.
  • Generalized Kohn-Sham caveat: Scaling constraints involving τ require care for Skala because its meta-GGA features are evaluated within the generalized Kohn-Sham scheme rather than pure Kohn-Sham.Only Ts + Exc is well defined in generalized Kohn-Sham meta-GGA calculations, so the self-consistent τ used for posterior checks must be interpreted accordingly.
  • Emergent kinetic-energy behavior: Skala learns T^GKS_c[ρ] to remain positive and close to the pure-Kohn-Sham value, rather than compensating higher kinetic energy with a more negative exchange-correlation energy.Training on MSR-ACC/TAE25 alone can allow the generalized-Kohn-Sham correlation kinetic energy to become negative, although accurate energies may still result.
  • Emergent kinetic-energy behavior: The positive, near-pure-Kohn-Sham T^GKS_c[ρ] behavior is consistent across 5 different random seeds, although the bound T^GKS_c[ρ] ≤ T_c[ρ] assumes identical densities.The densities used here are not exactly the same, so the inequality’s stated condition is not met exactly.

E.5 Ablation of the performance on geometry optimization · Supplementary References

Skala-1.1 performs accurately on geometry optimization tasks, including when reference molecular geometries are unavailable and geometries must be optimized with the functional. Ablations show that training on vibrationally distorted molecular geometries is important for achieving this accuracy.

  • E.5 Ablation of the performance on geometry optimization: The GEO metric measures the additional error incurred when a functional optimizes molecular geometries instead of evaluating energies only on reference geometries.It is defined as the energy difference between a functional evaluated at the reference geometry and at its own optimal geometry.
  • E.5 Ablation of the performance on geometry optimization: Skala-1.1 achieves W4-11-GEOM performance according to the GEO metric that is on par with the reported comparison.The passage states this result for the W4-11-GEOM test set and Fig. 15a.
  • E.5 Ablation of the performance on geometry optimization: Excluding TAEs of vibrationally distorted molecular geometries causes bond-length, bond-angle, and GEO performance to hover around revPBE.The ablation compares training data with and without MSR-ACC/Distortion.
  • E.5 Ablation of the performance on geometry optimization: Adding TAEs of vibrationally distorted molecular geometries boosts geometry-optimization performance from approximately revPBE-GGA levels to Skala-1.1’s current level.This supports the conclusion that distorted molecular geometries are important training data for accurate geometry optimization.
  • E.5 Ablation of the performance on geometry optimization: The results indicate that Skala-1.1 should perform well when reference molecular geometries are unavailable and the functional must optimize geometries.This conclusion follows from the geometry-optimization results and the distortion-data ablation.
  • Supplementary References: The supplementary geometry analyses compare optimized geometries across functionals and Skala checkpoints using the GEO metric on W4-11-GEOM.The figure description identifies GEO as reference 259 and W4-11-GEOM as reference 220.
Loading 2506.14665v6…