Source-linked AI summary
Deep-Learning Density Functional Theory Hamiltonian for Efficient ab initio Electronic-Structure Calculation
He Li, Zun Wang, Nianlong Zou, Meng Ye, Runzhang Xu, Xiaoxun Gong, Wenhui Duan, Yong Xu
TL;DR
DFT provides first-principles electronic-structure calculations but is computationally demanding, especially for large or twisted-material systems. DeepH learns DFT Hamiltonians with a locality-based, gauge- or rotation-covariant deep-neural-network framework, achieving high accuracy, efficiency, and transferability while supporting large-scale material studies. Its current training scope is limited when chemical environments vary substantially.
Problem
DFT electronic-structure calculations are computationally demanding for large systems, while twisted van der Waals materials require accurate methods beyond small-cell ab initio calculations.
Method
DeepH represents the DFT Hamiltonian using localized orbitals, locality, coordinate and basis transformations, and a message passing neural network.
Results
DeepH generally achieves high accuracy, efficiency, and transferability across material systems and physical properties, including twisted van der Waals materials.
Takeaways & Limitations
The framework extends first-principles research toward large-scale systems and enables investigation of twisted van der Waals materials.
Takeaways & Limitations
DeepH is currently applicable to unseen materials with chemical bonding environments close to those represented in its training dataset.
Abstract
from arXiv · showhide
The marriage of density functional theory (DFT) and deep learning methods has the potential to revolutionize modern computational materials science. Here we develop a deep neural network approach to represent DFT Hamiltonian (DeepH) of crystalline materials, aiming to bypass the computationally demanding self-consistent field iterations of DFT and substantially improve the efficiency of ab initio electronic-structure calculations. A general framework is proposed to deal with the large dimensionality and gauge (or rotation) covariance of DFT Hamiltonian matrix by virtue of locality and is realized by the message passing neural network for deep learning. High accuracy, high efficiency and good transferability of the DeepH method are generally demonstrated for various kinds of material systems and physical properties. The method provides a solution to the accuracy-efficiency dilemma of DFT and opens opportunities to explore large-scale material systems, as evidenced by a promising application to study twisted van der Waals materials.
INTRODUCTION
DFT is valuable for first-principles materials research but remains computationally demanding for systems exceeding thousands of atoms. Deep learning motivates learning the DFT Hamiltonian directly, despite its large dimensionality and gauge covariance.
- DFT has become indispensable across physics, materials science, chemistry, and biology, while deep learning has expanded scientific computation.
- Deep-learning potentials already learn interatomic interactions or potential energies from DFT for efficient molecular dynamics.
- The DFT Hamiltonian is the fundamental learned quantity because single-particle properties such as charge density and band structure can be derived from it.
- Learning the DFT Hamiltonian is harder than learning scalar quantities because its matrix transforms covariantly under coordinate, basis, and gauge changes.
- DeepH proposes a message passing neural network framework that uses locality to address Hamiltonian dimensionality and gauge or rotation covariance.
Theoretical framework of DeepH
DeepH addresses the large dimensionality and rotational or gauge covariance of DFT Hamiltonians by exploiting locality and localized-orbital representations. Local coordinates make transformed Hamiltonian blocks rotation-invariant before converting them back to the global frame.
- Challenges: DFT Hamiltonian learning is challenging because the matrix can have infinite dimensionality and must respect permutation, translation, rotation, and gauge transformations.These covariance requirements make Hamiltonian learning more difficult than learning scalar quantities such as total energy.
- Locality: Locality or electronic nearsightedness allows DeepH to focus on local environments rather than the entire material system.Local physical properties do not respond to distant changes in external potential because of destructive interference between many-particle eigenstates.
- Locality: Localized orbitals produce a sparse DFT Hamiltonian because matrix elements vanish beyond a distance threshold.This basis choice is compatible with the locality and non-periodicity of the problem and with the local or semilocal Kohn-Sham potential.
- Covariance handling: Local coordinate transformations make Hamiltonian blocks H′_ij invariant under rotation, after which rotation or basis transformations recover the global H_ij.This strategy avoids learning covariant relations directly through data augmentation, which is difficult for atom pairs with varying orientations.
Neural network architecture of DeepH
DeepH uses a message passing neural network on crystal graphs to predict locally transformed DFT Hamiltonian blocks, incorporating atomic, distance, neighborhood, and orientation information. An LCMP layer addresses local-coordinate instability before the predicted blocks are rotated back to obtain H_ij for unseen structures.
- Crystal graph: DeepH represents atoms as graph vertices and atom pairs within cutoff R_C as edges, using edge embeddings to represent H′_ij.Self-loop edges account for intra-site couplings.
- Feature construction: Initial vertex features encode atomic numbers, while initial edge features encode interatomic distances expanded with Gaussian basis functions.The graph therefore begins with chemical identity and pairwise geometric information.
- Message passing: Message-passing layers update vertex and edge features by aggregating neighborhood information within R_C, with stacked layers incorporating increasingly distant chemical environments.The resulting features are used to learn the locally transformed Hamiltonian blocks.
- Local-coordinate stability: Minor local-structure changes can substantially alter local coordinate axes and make H′_ij considerably different, reducing prediction accuracy.DeepH addresses this instability with an LCMP layer that adds bond orientation relative to the local coordinate after several message-passing layers.
- Orientation encoding: Real spherical harmonics and bond orientations provide orientation information for the local-coordinate representation.The local-coordinate features and edge features are updated to capture the geometry needed for predicting H′_ij.
- Workflow: After training on DFT data, DeepH predicts Hamiltonians for unseen atomic structures, bypassing time-consuming DFT self-consistent calculations.The predicted H_ij is obtained from H′_ij through a rotation transformation, enabling efficient electronic-structure calculations.
Capability of DeepH
DeepH accurately predicts DFT Hamiltonians and derived electronic properties across diverse materials, while substantially reducing computational cost. Its tests cover multiple atomic types, curved structures, unseen configurations, and large supercells.
- Graphene: DeepH predicts DFT Hamiltonians for graphene with an average test MAE of 2.1 meV across 13×13 orbital combinations.Individual MAEs range from 0.4 meV to 8.5 meV.
- Graphene: 1.9 meV is the average Hamiltonian MAE on 2,000 unseen graphene configurations sampled from 100 K to 400 K.This evaluates generalization beyond the training configurations.
- Electronic properties: DOS and shift-current spectra predicted by DeepH agree satisfactorily with DFT for unseen graphene configurations.The DOS MAE is on the order of 0.1 in units of 10^-3 eV^-1 Å^-2.
- Multiple elements: DeepH maintains high accuracy for MoS2 and systems with multiple atomic types, with Mo-Mo, Mo-S, S-Mo, and S-S MAEs of 1.3, 1.0, 0.7, and 0.8 meV.Band structures, electric susceptibility, and shift-current conductivity match DFT self-consistent calculations.
- Curved structures: DeepH generalizes to curved carbon and MoS2 nanotubes, with carbon-nanotube Hamiltonian MAE decreasing below 3.5 meV as diameter increases.The MAE is insensitive to nanotube chirality.
- Efficiency: DeepH has linear computational scaling with system size and reduces MoS2 35 × 35 supercell computation time by three orders of magnitude relative to DFT.DFT computational time roughly grows cubically with system size.
Application to twisted van der Waals materials
DeepH transfers from non-twisted training structures to twisted van der Waals materials, accurately reproducing Hamiltonians and material properties at scales difficult for conventional DFT. It also handles strong spin–orbit coupling while retaining the efficiency benefits of replacing self-consistent iterations.
- Workflow: Twisted-material training data can be generated from non-twisted structures, avoiding training over varying twist angles.The workflow trains DeepH on DFT data from relatively small, randomly perturbed supercells before applying it to new twisted structures.
- Twisted bilayer graphene: DeepH predicts twisted bilayer graphene properties accurately across varying twist angles after training at zero twist.Testing Moiré-twisted supercells reaches approximately 1,000 atoms with sub-meV average Hamiltonian MAE.
- Twisted bilayer graphene: DeepH reproduces the flat bands near the Fermi level in magic-angle twisted bilayer graphene.The magic-angle structure contains 11,164 atoms per supercell and its band structure matches a DFT benchmark.
- Strong spin–orbit coupling: DeepH achieves prediction accuracy comparable to twisted bilayer graphene for twisted bilayer bismuthenes with strong spin–orbit coupling.The method handles complex Hamiltonian values and spin degrees of freedom in the rotation transformation.
- Efficiency and scope: Replacing DFT self-consistent iterations with DeepH considerably reduces computational time for twisted materials, including magic-angle structures.Compared with empirical tight-binding and continuum models, DeepH is slightly less efficient but has better accuracy and transferability.
Wide applicability of DeepH
DeepH is presented as broadly applicable across material dimensionalities, geometries, compositions, and configuration spaces. Its locality-based treatment of covariance enables comparable accuracy to tensor-product methods with lower computational cost and fewer parameters.
- Material scope: DeepH demonstrates accuracy, efficiency, and transferability across quasi-1D and 2D materials with multiple elements, curved geometry, or Moiré twist.The authors also state that the method can be applied to other space dimensions.
- Generalization: DeepH’s PCA analyses examine generalization between monolayer sheets and nanotubes and between non-twisted and twisted bilayers.These analyses are used to assess output-feature behavior in geometrically different structures.
- Covariance and efficiency: DeepH achieves comparable accuracy to tensor-product covariant neural networks with much less computation time and fewer parameters on molecule datasets.Its approach performs the basis transformation once before training and uses rotation-invariant neural networks with local coordinates.
- Configuration space: DeepH generalizes to larger configuration spaces including 3D carbon allotropes and quasi-0D molecules without increased MAE relative to graphene.A unified neural network is used for graphite and diamond.
- Generalization: Experiments report Hamiltonian MAEs on the order of sub meV for additional systems, supporting applicability across a large configuration space.The authors identify these experiments as evidence for likely applicability to systems spanning a large configuration space.
DISCUSSION
DeepH maps material structures to physical properties and extends first-principles studies toward large-scale and twisted van der Waals systems. Its current transferability is limited when chemical environments differ strongly, and such cases require more deliberately constructed training data.
- Contribution: DeepH builds a mapping from material structures to physical properties using a deep-neural-network representation of the DFT Hamiltonian.The framework is intended to extend first-principles research to large-scale material systems.
- Limitation: The trained model is currently applied to unseen materials only when their chemical bonding environments are close to those in the dataset.Strongly varying chemical environments require manually designed datasets to improve training efficiency.
- Future work: Automatic dataset construction and on-the-fly training optimization are identified as future directions.These directions address the need to improve training efficiency for strongly varying chemical environments.
- Extensions: DeepH can in principle describe disordered, defective, and interfacial systems, but these cases demand more training data for varying chemical environments.Non-periodic large-scale systems and several advanced DFT functionals are described as possible extensions, with hybrid functionals requiring a larger cutoff radius.
Dataset preparation
The study prepares training datasets from ab initio molecular-dynamics configurations and perturbed bilayer structures for several crystalline and Moiré materials. These datasets use specified supercells, computational settings, interlayer spacings, and structure counts.
- Monolayer datasets: 6 × 6 graphene and 5 × 5 MoS2 monolayer supercells provide random configurations generated by ab initio molecular-dynamics calculations.The simulations use VASP, projector-augmented wave pseudopotentials, PBE, a 450 eV plane-wave cutoff, and Γ-point sampling.
- Moiré-twisted datasets: 3.35 Å for TBG and 3.20 Å for TBB are the interlayer spacings used for training and Moiré-twisted supercells.These spacings come from fully relaxed bilayer unit cells with the most energetically favorable stacking.
- Moiré-twisted datasets: 300 TBG and 576 TBB shifted and perturbed supercell structures are included in the respective datasets.The structures combine layer translations with random perturbations at atomic sites.
- Hamiltonian generation: OpenMX calculations generate DFT Hamiltonians using PBE, norm-conserving pseudopotentials, and material-specific pseudo-atomic localized orbitals.Graphene, CNTs, and TBG use C6.0-s2p2d1 orbitals with 13 atomic-like basis functions and RC = 6.0 Bohr; MoS2 systems use separate Mo and S orbital sets.
Physical properties derived from DFT Hamiltonian
The DFT Hamiltonian and overlap matrix yield eigenvalues and eigenstates through a generalized eigenvalue problem, enabling calculation of electronic and response properties. The framework also adapts susceptibility and conductivity expressions for low-dimensional systems.
- Electronic structure: Hamiltonian and overlap matrix elements in a non-orthogonal atomic-orbital basis are Fourier transformed and solved as a generalized eigenvalue problem to obtain bands and eigenstates.The overlap matrix is computed directly from basis-function inner products rather than learned by a neural network.
- Electronic structure: A few eigenvalues of the large-scale sparse Hamiltonian for Moiré-twisted materials are computed with the ARPACK library.The sparse Hamiltonian is obtained from the DeepH method.
- Response properties: The framework calculates 3D electric susceptibility and shift-current conductivity as functions of light frequency.The response expressions use band energies, occupations, Berry connections, and their derivatives derived from the DFT Hamiltonian.
- Low-dimensional responses: For low-dimensional systems, response functions are redefined to remove vacuum-layer effects from supercell calculations.The study considers 2D MoS2 susceptibility, 1D CNT susceptibility, and 2D graphene sheet conductivity using supercell geometry factors.
- Low-dimensional responses: χ describes electric susceptibility along the periodic direction in the low-dimensional response formulation.
Details on training neural network
DeepH uses message-passing neural networks to update vertex and edge features and represent Hamiltonian matrix elements. The implementation combines multiple message-passing layers, localized graph cutoffs, and alternative output strategies for different materials.
- Network architecture: The neural networks update both vertex features and edge features within the message-passing architecture.Vertex and edge transformations use neural-network components described for the model.
- Network architecture: The vertex network uses inputs, weights, biases, element-wise multiplication, sigmoid activation, and softplus activation.The supplied formulation identifies these operations and their associated dimensions or meanings.
- Network architecture: The edge network is a fully connected neural network with one hidden layer and a SiLU activation function.
- Training configuration: Five message-passing layers and one LCMP layer give the model 471409 + 129 × Nout parameters.The number of selected orbital pairs Nout is a model hyperparameter; graph construction uses the corresponding localized-orbital cutoff radius.
- Training configuration: The model uses 64-dimensional elemental embeddings and vertex features, while Adam learning rates decrease from 1×10^-3 to 2×10^-4 and finally 4×10^-5.The implementation uses PyTorch-Geometric.
- Output representation: Separate MPNN models represent different orbital pairs for graphene and TBG, whereas single models output multidimensional Hamiltonian blocks for MoS2 and TBB.The latter strategy is used to improve efficiency.
COMPETING INTERESTS
The figures outline DeepH’s locality-based representation of DFT Hamiltonians, its message-passing architecture, and evaluations across graphene, MoS2, nanotubes, and twisted materials.
- FIG. 1: DeepH learns DFT Hamiltonians using locality, with localized-basis matrix elements restricted to neighboring atoms within a cutoff radius.The framework also accounts for covariance under unitary or rotation transformations.
- FIG. 2: The DeepH model uses crystal graphs and message passing, including a local-coordinate message-passing layer with additional orientation information.The architecture contains L layers and applies different local coordinates to atom pairs.
- FIG. 3: For graphene, performance is evaluated through orbital-resolved Hamiltonian MAE and nearest-neighbor matrix-element distributions.The figure also compares generalization across unseen structures using density of states and shift current conductivity.
- FIG. 4: For monolayer MoS2, the evaluation compares orbital-resolved MAE, DFT and DeepH band structures, and real and imaginary electric susceptibility.The susceptibility and band-structure comparisons use a 5 × 5 supercell representative of median generalization error.
- FIG. 5: DeepH is tested for transfer from flat sheets to curved graphene and MoS2 nanotubes by comparing their band structures with DFT.The examples include zigzag (25, 0) carbon and (50, 0) MoS2 nanotubes.
- FIG. 6: For twisted materials, the workflow trains on small non-twisted structures, predicts arbitrary twist angles, compares band structures, and contrasts Hamiltonian-construction times with DFT.The examples cover twisted bilayer graphene and twisted bilayer bismuthene across multiple twist angles.