Source-linked AI summary
Neural Network Potentials: A Concise Overview of Methods
Emir Kocer, Tsz Wai Ko, Jörg Behler
TL;DR
Atomistic simulations need efficient potentials that retain the accuracy of high-level electronic-structure calculations across diverse systems. This review surveys neural-network potentials, organizing them by descriptors, dimensionality, and long-range physics, and discusses their capabilities and open challenges. NNPs now support large-scale simulations, but transferability beyond training data remains a central limitation.
Problem
MLPs must represent complex potential energy surfaces efficiently while addressing locality, descriptor invariance, long-range interactions, and non-local charge transfer across diverse systems.
Method
The review classifies neural-network potentials by descriptor type, system dimensionality, long-range electrostatics, and treatment of global structure and charge distribution.
Results
NNPs provide accurate, analytic potentials for large-scale atomistic simulations, with typical HDNNP errors around 1.0 meV for energies and 0.1 eV/Å for forces.
Takeaways & Limitations
NNPs are useful for large-scale molecular dynamics across chemistry, molecular biology, and materials science, including systems with diverse bonding types.
Takeaways & Limitations
NNPs have restricted transferability beyond the structural diversity of their underlying training data.
Abstract
from arXiv · showhide
In the past two decades, machine learning potentials (MLP) have reached a level of maturity that now enables applications to large-scale atomistic simulations of a wide range of systems in chemistry, physics and materials science. Different machine learning algorithms have been used with great success in the construction of these MLPs. In this review, we discuss an important group of MLPs relying on artificial neural networks to establish a mapping from the atomic structure to the potential energy. In spite of this common feature, there are important conceptual differences, which concern the dimensionality of the systems, the inclusion of long-range electrostatic interactions, global phenomena like non-local charge transfer, and the type of descriptor used to represent the atomic structure, which can either be predefined or learnable. A concise overview is given along with a discussion of the open challenges in the field.
I. INTRODUCTION AND SCOPE OF THIS REVIEW
Machine learning potentials use reference electronic-structure data to represent potential energy surfaces, enabling faster atomistic simulations. This review focuses on neural-network potentials and organizes them by system dimensionality, descriptor type, and treatment of long-range effects.
- MLPs learn multidimensional potential energy surfaces from high-level electronic-structure reference data for atomistic simulation.The potential energy surface contains information about structures, forces, transition states, barriers, and vibrations.
- Analytic MLPs can run molecular dynamics many orders of magnitude faster than electronic-structure calculations without significant accuracy loss.
- MLPs provide flexible, general functional forms for diverse bonding types and consistent energy, force, and stress calculations.
- Their main disadvantage is limited transferability beyond the structural diversity represented in training data, making careful validation essential.
- The review surveys neural-network potentials with continuous potential energy surfaces for arbitrary atomic positions, while excluding many property-prediction approaches.
- It classifies NNPs using descriptor type, system dimensionality, and inclusion of long-range electrostatic interactions, while placing them in historical context.
III. THE SECOND GENERATION: LOCAL METHODS FOR HIGH-DIMENSIONAL SYSTEMS
Early MLPs were restricted by low dimensionality and by the difficulty of constructing invariant descriptors for variable-size, high-dimensional systems. Second-generation methods addressed these limitations through local atomic-energy decompositions and a broad range of descriptors.
- First-generation NNPs explicitly represented only a few degrees of freedom, limiting applications mainly to small molecules and other low-dimensional systems.
- High-dimensional applications required descriptors invariant to translation, rotation, and permutation of chemically equivalent atoms.
- A single global neural network also prevented systems with variable atom counts because input dimensionality is fixed after training.
- Second-generation MLPs use many available high-dimensional descriptors and are applicable in principle to systems of arbitrary size.
- Second-generation NNPs divide into predefined-descriptor and learnable-descriptor classes, with predefined descriptors using functional forms and a small parameter set.
B. Predefined Descriptors
Predefined-descriptor NNPs transform Cartesian coordinates into representations of local atomic environments and predict atomic or pair contributions to short-ranged energy. Their descriptors encode radial and angular neighbor structure while using cutoff functions to impose locality.
- High-Dimensional Neural Network Potentials: HDNNPs use local atomic environments and cutoff-based descriptors to represent atomic contributions while preserving required symmetries.
- High-Dimensional Neural Network Potentials: The cutoff function smoothly decays to zero in value and slope at radius Rc, with Rij denoting the distance between atoms i and j.
- Second-generation predefined-descriptor NNPs construct short-ranged total energy as sums of atomic or pair neural-network outputs.
- High-Dimensional Neural Network Potentials: Radial ACSFs characterize neighbor distributions using Gaussian-shaped functions, while angular ACSFs probe the angles formed by neighboring atoms.
- High-Dimensional Neural Network Potentials: For multi-element systems, separate radial pair and angular triple element combinations cause a combinatorial increase in symmetry functions.
- High-Dimensional Neural Network Potentials: Typical HDNNP errors are around 1.0 meV for energies and 0.1 eV/Å for forces across systems including bulk materials, solutions, interfaces, and organic molecules.
2. Pair-Based HDNNPs
Pair-based HDNNPs express total energy through environment-dependent pair or bond contributions. These approaches can preserve symmetry, target specific interactions, or improve interpretability, but pair-based evaluation can increase computational cost.
- Pair-based HDNNPs: Pair-based HDNNPs replace atoms as central structural entities with atom pairs and use environment-dependent pair energies.Pair symmetry functions characterize both atoms’ combined environments up to a cutoff while preserving translational, rotational, and permutational symmetry.
- Pair-based HDNNPs: Pair symmetry functions produce accuracy comparable to atom-based approaches but require more neural-network evaluations because systems contain more pairs than atoms.
- Bond-based models: BIM-NN represents energy as bond contributions using bond lengths and information about directly connected bonds within element-specific distance cutoffs.The resulting pair energies are interpreted as chemical bond energies.
- Interaction-focused models: AP-Net focuses on non-covalent interactions by representing physically meaningful interaction energies derived from symmetry adapted perturbation theory.A single neural network can represent all element combinations for a given physical interaction type.
3. Deep Potential Molecular Dynamics
Deep Potential Molecular Dynamics uses local atomic frames and jointly processed neighbor descriptors to predict environment-dependent atomic energies. Related descriptor approaches range from predefined structural fingerprints to learned representations and embedding-inspired densities.
- DeepMD: DeepMD writes total energy as a sum of environment-dependent atomic energies and describes neighbors in a local frame based on the two closest atoms.Neighbors are sorted by element and inverse distance before descriptor construction.
- DeepMD: DeepMD descriptors avoid tunable hyperparameters and depend only on atomic positions, reducing potential complexity.
- DeepMD: DeepMD’s simple descriptors lack explicit many-body information, which must be established by the atomic neural networks.
- DeepMD: DeepMD has force discontinuities at the cutoff radius because it applies no smooth cutoff function; DeepMD-SE addresses this with scalar weight functions.
- EANN: EANN replaces EAM’s scalar surrounding-atom density with a vector of GTO-based densities whose expansion coefficients are optimized during training.
- Descriptor strategies: Predefined descriptors provide local structural fingerprints as inputs to atomic neural networks, whereas learned descriptors replace static representations with dynamically learned features.
2. Deep Tensor Neural Networks (DTNN)
DTNN is a global message-passing neural network for small molecules that refines atomic features using all pairwise distances before predicting atomic energies. Its accuracy has been demonstrated across compositional, configurational, and conformational spaces.
- DTNN architecture: DTNN starts from nuclear charges and a complete distance matrix, expands pairwise distances in a Gaussian basis, and assigns each atom a charge-specific feature vector.
- DTNN architecture: DTNN performs T = 3 global refinement steps, after which accuracy saturates for the investigated systems.
- DTNN architecture: Tensor-layer message vectors nonlinearly couple neighboring atomic features with corresponding distance vectors before atomic neural networks yield atomic energies.
- Applications and scope: DTNN accuracy was demonstrated for optimized-molecule atomization energies, isomer configurations, and molecular-dynamics trajectories of small molecules in vacuum.
- Applications and scope: Using complete interatomic distances makes DTNN a global method for small molecules, although spatial cutoffs could extend it toward second-generation potentials.
3. SchNet
SchNet improves on DTNN with continuous-filter convolutional layers that represent continuous atomic environments. Its refined atomic features are converted into atomic energy contributions and summed into the total energy, while HIP-NN adds hierarchical many-body decomposition.
- SchNet: SchNet uses continuous-filter convolutional layers to represent atomic environments with continuous positional information.The layers generalize discrete convolutional networks used for image grids, where atomic positions are continuous quantities.
- SchNet: After interaction-block refinement, SchNet feeds atomic feature vectors into atomic neural networks whose energy contributions are added to obtain total energy.
- HIP-NN: HIP-NN combines message passing with a conventional many-body expansion by decomposing atomic energies into contributions of order n.
- HIP-NN: HIP-NN learns the many-body contributions simultaneously through one hierarchical neural network with multiple layers and intermediate interaction blocks.
- HIP-NN: HIP-NN uses nuclear charges and pairwise distances below a predefined cutoff, then terminates the expansion after the n = 2 term.
- HIP-NN: The magnitudes of HIP-NN’s different energy terms can estimate prediction uncertainty because contributions are assumed to decay with increasing order.
5. AIMNet
AIMNet uses learnable atomic descriptors and message passing to predict atomic energies and properties, but its treatment of long-range electrostatics remains limited by the interaction range of those updates.
- Architecture: AIMNet constructs atomic feature vectors from modified ACSFs and nuclear-charge embeddings, whose dimensionality is independent of the number of elements.The ACSF-based geometric description is combined with initial feature vectors containing nuclear-charge information.
- Architecture: AIMNet can apply an embedding layer immediately because its initial feature vectors already encode atomic environments.Unlike message-passing networks requiring a preparatory update, AIMNet starts with environment information from ACSFs.
- Interaction range: Repeated interaction blocks extend the effective cutoff of atomic interactions by updating feature vectors through multiple message-passing steps.The AIM layer is typically calculated three times, corresponding to two message-passing updates.
- Predicted properties: AIMNet predicts environment-dependent atomic energies and can also provide atomic partial charges redistributed within the message-passing interaction range.AIMNet-ME additionally uses a message-passing stack for spins and charges to treat systems with different global charge states.
- Limitation: Long-range electrostatics are not explicitly included beyond the distance covered by the message-passing steps.This limits AIMNet's treatment of interactions extending beyond its effective message-passing range.
- Context: Third-generation NNPs explicitly compute untruncated Coulomb interactions, but their added computational costs yield small accuracy improvements in many screened condensed systems.These methods are therefore rarely used despite incorporating long-range electrostatics.
V. THE FOURTH GENERATION: NON-LOCAL INTERACTIONS
Fourth-generation NNPs address non-local or global electronic dependencies that local-charge models cannot represent reliably. The review distinguishes these dependencies from untruncated but locally parameterized long-range interactions.
- Motivation: Local-charge assumptions fail when partial charges depend on distant functional groups, defects, doping, ion substitution, or global charge state.The first three NNP generations implicitly assume a single fixed total charge and cannot reliably represent such potential energy surfaces.
- Definition: Fourth-generation NNPs are defined for high-dimensional systems that capture non-local or global dependencies in the electronic structure.These dependencies can include non-local charge transfer and changes in global charge state.
- Terminology: The review uses long-range interactions for untruncated electrostatics or van der Waals interactions depending on local properties, whereas non-local interactions require fourth-generation NNPs.The terminology is not consistent across the literature, so the review makes this distinction explicitly.
1. Charge Equilibration Neural Network Technique
CENT combines local neural-network electronegativities with charge equilibration to obtain globally dependent atomic charges and minimize the electrostatic energy. The section places CENT among fourth-generation methods while noting a bonding-specific scope boundary.
- Charge equilibration: Charge equilibration redistributes electrons across the whole system while minimizing the electrostatic energy under a total-charge constraint.The minimization produces equilibrated atomic partial charges by solving a set of linear equations.
- CENT: CENT expresses atomic electronegativities as neural-network functions of local environments described by ACSFs.The neural networks provide electronegativities χi for the charge-equilibration model.
- Training: CENT trains electronegativity-network weights by minimizing the error of its energy against DFT reference energies.The charge-equilibration solution supplies the globally dependent charges used in the energy expression.
- Scope: The CENT method works best for systems with predominantly ionic bonding.
- BpopNN: BpopNN determines atomic populations self-consistently by minimizing total energy for a chosen global charge.Its energy includes atomic neural-network contributions, intra-atomic energy, and electrostatic energy.
- BpopNN: In BpopNN, atomic neural networks use local environments and atomic populations through modified SOAP descriptors, while charges also enter intra-atomic and electrostatic terms.Training uses constrained-DFT populations, energies, and forces across charge distributions and states.
3. Fourth-Generation High-Dimensional Neural Network Potentials
4G-HDNNPs combine charge-equilibration-based non-local electronic information and electrostatics with the accurate local bonding description of second-generation HDNNPs. They are presented as broadly applicable, while descriptor scalability and training-data efficiency remain open challenges.
- 3. Fourth-Generation High-Dimensional Neural Network Potentials: 4G-HDNNPs combine CENT's non-local charge transfer and global electronic dependencies with second-generation HDNNPs' accurate local bonding description.The method also includes the resulting electrostatic interactions.
- Energy model: Their short-range energy sums atomic contributions augmented by the respective partial charges as additional inputs encoding local electronic structure.Short-range atomic networks are trained after the electrostatic networks in a second step.
- Applicability: 4G-HDNNPs apply to systems ranging from organic molecules to ionic solids and can simultaneously describe different global charge states.
- Descriptors: Predefined and learnable descriptors are both suitable for high-quality potentials, but predefined descriptors face combinatorial growth as the number of elements increases.The generality of combined descriptors, combined atomic networks, and combined feature vectors remains to be explored.
- Training data: Reference datasets must be small enough to limit expensive electronic-structure calculations while covering broad structural variation for transferability.
- Data efficiency: Delta-learning uses a baseline potential and learns the difference to a higher-level potential when that difference is smooth.
- Physical information: Physical information such as electrostatics, spin moments, and external electric fields is increasingly incorporated to improve or extend neural-network potentials.The review also identifies emerging message-passing approaches for non-local effects and electronic-structure-informed representations.
- Outlook: Restricted transferability beyond the training set remains a main limitation of NNPs and MLPs generally.The review anticipates that future potentials may address this through physical knowledge while retaining broad applicability.
DISCLOSURE STATEMENT
The authors disclose no affiliations, memberships, funding, or financial holdings that might be perceived as affecting this review’s objectivity.
- The authors report no affiliations, memberships, funding, or financial holdings that might be perceived as affecting the review’s objectivity.