Source-linked AI summary
Coupled-cluster molecular properties across the main group that extrapolate beyond training size
Wenhao He, Xu Chen, Noah Song, Haowei Xu, Tim S. Hindges, Bohan Li, Zihan Lin, Yu Yao, Avetik R. Harutyunyan, Fang Liu, Yao Wang, Hao Tang, Ju Li
TL;DR
Coupled-cluster accuracy is costly, while affordable density-functional calculations are systematically biased. This paper introduces MEHnet-MG, which corrects a single DFT calculation through a predicted Hamiltonian and delivers coupled-cluster-quality molecular properties, including size-extrapolating oligothiophene predictions.
Problem
Molecular-property predictions face a trade-off between coupled-cluster accuracy and the lower cost but systematic bias of density-functional theory.
Method
MEHnet-MG uses an equivariant network to correct DFT-derived electronic-structure objects and derives seven molecular properties from the corrected Hamiltonian across nine main-group elements.
Results
Across the main group, one model reaches coupled-cluster accuracy for a broad property suite; it tracks finite-field CCSD polarizability to ∼2% out to 44 atoms.
Takeaways & Limitations
The method enables high-throughput coupled-cluster-quality predictions for six properties across nine main-group elements at a cost dominated by one inexpensive DFT calculation.
Takeaways & Limitations
Optical-gap accuracy is limited by the EOM-CCSD/cc-pVDZ labels, whose basis, method, and growing double-excitation character make the gap label non-gold-standard.
Abstract
from arXiv · showhide
Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. We resolve this trade-off with a single equivariant network, MEHnet-MG, that predicts an effective one-electron Hamiltonian from one inexpensive B3LYP/def2-SVP calculation and derives a broad suite of properties from it (energy, optical gap, dipole, quadrupole, polarizability, Mulliken atomic charges, and Mayer bond orders) at coupled-cluster accuracy across nine main-group elements, including the under-served phosphorus, sulfur, and chlorine chemistries. The model is trained on a new in-house dataset of multi-property labels computed at the CCSD(T) level for all nine elements. On a held-out test set, it reduces the error of every property by a factor of 3.8 to 230 relative to semi-local, hybrid, and double-hybrid DFT (referenced to composite CCSD(T)/cc-pVTZ; Methods), while adding only ~25 ms wall time per molecule, delivering coupled-cluster-quality predictions at the cost of a single DFT calculation. Critically, deriving every property from a predicted Hamiltonian rather than pooling per-atom features builds the correct size-scaling into the model architecture: on pi-conjugated oligothiophenes it matches finite-field CCSD polarizability and the EOM-CCSD optical gap to ~2% at the largest sizes where those references remain affordable (44 and 37 atoms, where a single CCSD field point already costs ~500x the model's entire inference) and extrapolates the corrected trends to 58-atom chains, a regime where pooling-based architectures fail by construction. Accurate extrapolation is therefore set by the model's inductive bias rather than by the training data.
Results
MEHnet-MG derives seven molecular properties from a corrected effective Hamiltonian and screening matrix, achieving coupled-cluster-level accuracy across nine main-group elements at near-DFT cost. Its Hamiltonian-based read-out also captures nonlocal size scaling, enabling accurate extrapolation to longer π-conjugated chains where pooling-based architectures fail.
- Unified property prediction: Seven properties for nine main-group elements emerge from one corrected effective Hamiltonian and screening matrix rather than separate property predictors.Diagonalizing the corrected Hamiltonian supplies orbital energies, density-derived observables, charges, and bond orders; the corrected screening matrix supplies polarizability.
- Accuracy across chemistry: MEHnet-MG leads every property in the table and generalizes uniformly across chemistry, with no degradation for fluorine and silicon relative to carbon, hydrogen, nitrogen, and oxygen.The model’s single predicted-Hamiltonian read-out reports the complete suite at one accuracy level, while B3LYP is best only for energy and the double hybrid’s orbital gap remains poor for the EOM-CCSD optical gap.
- Experimental validation: 0.078 Debye is the model’s mean absolute error against measured dipole magnitudes across twelve molecules spanning all nine elements.Canonical CCSD(T)/def2-TZVPP achieves 0.046 Debye against the same experiments, while the model improves over its B3LYP/def2-SVP baseline at the baseline’s cost.
- Extrapolation to conjugated chains: ∼2% is the signed isotropic polarizability error against finite-field CCSD/cc-pVDZ through T6, while MEHnet-MG continues the corrected trend to T8.The isotropic polarizability rises ∼12-fold from T1 to T8, demonstrating that the model tracks a super-linear response beyond its training chemistry.
- Extrapolation to conjugated chains: 0.04 eV is the model’s EOM-CCSD/cc-pVDZ optical-gap agreement at T2–T5, despite the gap closing and saturating with conjugation length.The T1 monomer is an exception, deviating by −0.37 eV; its lowest EOM root is nevertheless confirmed as the targeted HOMO→LUMO single excitation.
- Inductive bias and size scaling: Pooling local contributions cannot represent super-linear collective response, whereas diagonalizing the predicted Hamiltonian produces frontier orbitals that delocalize over the conjugation length.The predicted Hamiltonian’s frontier quantities continue evolving beyond the 6 Å cutoff, while pooled read-outs are structurally limited to extensive or intensive scaling.
Discussion
MEHnet-MG combines coupled-cluster-quality molecular-property prediction with the cost of a single DFT calculation across nine main-group elements. Its Hamiltonian-based architecture enables correct size scaling and extrapolation, while optical-gap labels and dataset splits remain important limitations.
- Accuracy–cost trade-off: A single equivariant model reaches coupled-cluster accuracy for broad molecular properties across the main group at the cost of one DFT single point.The method covers nine main-group elements and places accuracy and cost at a favorable trade-off.
- Architectural size scaling: Pooling per-atom contributions cannot represent super-linear collective responses, making correct size scaling an architectural rather than learnable property.Such pooling yields only strictly extensive or strictly intensive scaling.
- Oligothiophene extrapolation: ∼2% polarizability error persists to the 44-atom CCSD limit, while the model continues the corrected super-linear trend beyond that limit.One CCSD field point costs 5.7 h, approximately 500× the model’s entire inference.
- Practical scope and limitation: The method enables high-throughput coupled-cluster-quality prediction of six ground-state and response properties for nine main-group elements, including phosphorus, sulfur, and chlorine.The cost is dominated by a single cheap DFT call; the optical gap is an exception because it faithfully follows its EOM-CCSD/cc-pVDZ label and its offset from experiment.
- Optical-gap limitation and roadmap: The optical gap is limited by EOM-CCSD/cc-pVDZ labels that omit triples, lack diffuse functions, and increasingly encounter double-excitation character with chain length.A proposed roadmap adds diffuse functions, uses the fundamental gap IP−EA, and develops out-of-distribution conjugated-molecule holdouts.
- Evaluation limitations: Shuffled dataset-index splits do not ensure chemically distant test molecules, so scaffold-split benchmarks would more conservatively assess transfer to unseen chemistry.The comparison with other machine-learning models is also necessarily heterogeneous.
Methods
MEHnet-MG is trained on a 41,939-molecule, nine-element main-group dataset with coupled-cluster property labels and predicts an effective Hamiltonian from a B3LYP/def2-SVP calculation. Quantum-mechanical readout from that Hamiltonian provides the target properties while preserving system-size dependence, with diabatic level tracking stabilizing training.
- Dataset: 41,939 closed-shell neutral molecules spanning H, C, N, O, F, Si, P, S, and Cl form the dataset, split into 40,985 training, 968 validation, and 959 test molecules.Structures contain 2–24 atoms and at most nine heavy atoms after filtering 44,412 PubChem structures.
- Reference labels: Composite CCSD(T)/cc-pVTZ labels cover energy, dipole, quadrupole, Mulliken charges, and Mayer bond orders, while finite-field CCSD/cc-pVDZ labels polarizability and EOM-CCSD/cc-pVDZ labels the lowest singlet vertical gap.The composite coupled-cluster estimate combines canonical CCSD(T)/cc-pVDZ with a DLPNO-CCSD(T) basis-set correction to cc-pVTZ.
- Property readout: Every property is evaluated from the predicted Hamiltonian and its eigenstates using the defining quantum-mechanical operator rather than a separate learned property head.Energy, the HOMO–LUMO gap, multipoles, polarizability, Mulliken charges, and Mayer bond orders are all derived from the Hamiltonian or its density matrix.
- Hamiltonian model: The model predicts a correction to the B3LYP/def2-SVP Fock matrix, constructs H = HDFT+∆Hθ, and derives molecular orbitals and the density matrix by solving its eigenvalue equations.The network is SO(3)-equivariant and uses spherical-harmonic graph attention with lmax = mmax = 4 to represent third-row on-site Fock blocks.
- Training stability: Diabatic level tracking prevents orbital crossings from discontinuously reassigning the HOMO and destabilizing density-derived response-property gradients during training.The failure mode arises because learned Hamiltonian corrections shift the spectrum and strict energy ordering can change orbital assignments.