Source-linked AI summary
The Design Space of E(3)-Equivariant Atom-Centered Interatomic Potentials
Ilyes Batatia, Simon Batzner, Dávid Péter Kovács, Albert Musaelian, Gregor N. C. Simm, Ralf Drautz, Christoph Ortner, Boris Kozinsky, Gábor Csányi
TL;DR
Machine-learning interatomic potentials include diverse ACE and equivariant message-passing architectures whose design choices are difficult to compare systematically. The paper introduces Multi-ACE to unify them, studies NequIP choices through ablations, and proposes BOTNet, which maintains benchmark accuracy with a simpler body-ordered architecture.
Problem
The paper addresses the lack of a unified framework for comparing the expanding design space of ACE and equivariant message-passing interatomic potentials.
Method
Multi-ACE recasts ACE as layers in a multi-layer equivariant architecture and uses systematic NequIP ablations to guide a simplified BOTNet design.
Results
BOTNet and NequIP achieve similar accuracy on the reported benchmark tests, while message normalization decreases error by over 30% in the compared high-temperature setting.
Takeaways & Limitations
The framework identifies internal normalization and selected equivariant architectural elements as important design choices for accuracy, smoothness, and extrapolation.
Abstract
from arXiv · showhide
The rapid progress of machine learning interatomic potentials over the past couple of years produced a number of new architectures. Particularly notable among these are the Atomic Cluster Expansion (ACE), which unified many of the earlier ideas around atom density-based descriptors, and Neural Equivariant Interatomic Potentials (NequIP), a message passing neural network with equivariant features that showed state of the art accuracy. In this work, we construct a mathematical framework that unifies these models: ACE is generalised so that it can be recast as one layer of a multi-layer architecture. From another point of view, the linearised version of NequIP is understood as a particular sparsification of a much larger polynomial model. Our framework also provides a practical tool for systematically probing different choices in the unified design space. We demonstrate this by an ablation study of NequIP via a set of experiments looking at in- and out-of-domain accuracy and smooth extrapolation very far from the training data, and shed some light on which design choices are critical for achieving high accuracy. Finally, we present BOTNet (Body-Ordered-Tensor-Network), a much-simplified version of NequIP, which has an interpretable architecture and maintains accuracy on benchmark datasets.
I. Introduction
The paper positions ACE descriptors and equivariant message-passing networks within a shared atomistic modeling landscape. Multi-ACE uses ACE constructions in successive layers, enabling systematic comparison of architectural choices and motivating BOTNet.
- I. Introduction: ACE provides complete, body-ordered basis functions for atomic environments, encompassing several earlier symmetrized descriptors.Its basis uses spherical harmonics, radial functions, and controlled body order.
- I. Introduction: Equivariant networks retain features that transform under symmetry operations and construct invariant outputs only at the final stage.Equivariant tensor products couple features while preserving their prescribed transformation behavior.
- I. Introduction: Multi-ACE unifies ACE and message passing by using ACE in each network layer, creating a broad design space for interatomic potentials.The framework supports systematic probing of choices affecting model behavior.
- I. Introduction: Message-passing potentials represent atoms as cutoff graphs, update learned features through neighbor aggregation, and map atomic states to site energies.The graph encodes spatial relationships rather than only conventional chemical bonds.
- I. Introduction: In an MPNN, neighbor-derived messages are aggregated, combined with central-atom features, and passed through learned update and readout functions.Nonlinear updates can generate higher body-order dependencies from lower-order messages.
C. Equivariant messages
Equivariant messages are constructed by imposing rotational transformation laws on basis functions and features. The ACE construction begins with atom-pair basis functions, pools neighbors into permutation-invariant quantities, and extends them through products and symmetry-aware coupling.
- C. Equivariant messages: Equivariance requires messages to transform according to irreducible representations of the Euclidean symmetry group.A message with label L transforms through the corresponding Wigner-D representation under rotations.
- C. Equivariant messages: Products of pooled basis functions create higher body-order features, while explicit body-order expansions can be truncated when higher-order terms are sufficiently small.This provides a systematic approximation strategy for high-dimensional functions.
- C. Equivariant messages: Nonlinear activations generally prevent an MPNN from remaining explicitly body-ordered, whereas linear update and readout functions preserve that structure.The distinction motivates analyzing activation choices separately from equivariant feature construction.
- C. Equivariant messages: Pooling one-particle basis values over neighbors produces permutation-invariant atomic features that serve as the starting point for many-body constructions.The one-particle basis can be viewed as edge features in a graph model.
- C. Equivariant messages: ACE one-particle basis functions combine orthogonal radial functions with spherical harmonics and depend on central and neighboring atom states.Chemical attributes can be discrete labels or continuous learned embeddings.
B. Higher order basis functions
ACE generates high body-order features efficiently from pooled one-particle basis functions rather than explicit enumeration of neighbor tuples. Products of pooled functions provide a complete permutation-invariant basis with directly controlled correlation and body order.
- B. Higher order basis functions: The ACE density trick evaluates high body-order features without explicitly summing over triplets, quadruplets, or larger clusters.This enables systematic body-ordered expansions at low computational cost.
- B. Higher order basis functions: Summing one-particle basis values over neighbors creates the permutation-invariant A-basis, whose elements are two-body functions.These functions depend on all neighbor positions through decomposable sums.
- B. Higher order basis functions: Products of A-basis functions form a complete basis of permutation-invariant atomic-environment functions.The product basis is indexed by correlation order and coupled multi-indices.
- B. Higher order basis functions: A product of ν A-basis functions has correlation order ν and body order ν + 1 because the central atom is included.For example, ν = 3 produces four-body basis functions.
- B. Higher order basis functions: Linear ACE forms products over radial, angular, and chemical indices in the coupled channels, without uncoupled k channels.This distinguishes linear ACE from the channel organization used in the generalized construction.
C. Symmetrization of basis functions
Symmetrization converts permutation- and translation-invariant product bases into rotationally invariant or equivariant features. Multi-ACE then propagates these ACE-derived messages through layers using symmetry-compatible updates and a learned readout.
- C. Symmetrization of basis functions: Rotational symmetrization averages product-basis functions over O(3) to obtain invariant or equivariant basis functions.The invariant case is expressed through an integral over rotations.
- C. Symmetrization of basis functions: Equivariant basis functions are constructed to transform in the same representation as the features they expand.The framework applies to tensors in Cartesian or spherical coordinates.
- C. Symmetrization of basis functions: Rotation integrals can be evaluated as tensor contractions using Wigner-D products and generalized coupling coefficients.These operations generate spanning sets of equivariant features.
- C. Symmetrization of basis functions: The B functions span all many-body atomic-environment functions with a specified symmetry and can be linearly combined into atom-wise messages.A learned transformation mixes the channel index while retaining the imposed symmetry label.
- C. Symmetrization of basis functions: Multi-ACE feeds each layer’s outputs into the next layer, where learnable updates mix uncoupled message channels through block-diagonal equivariant weights.Only features with matching representations interact linearly.
- C. Symmetrization of basis functions: After the final layer, a learned linear or nonlinear readout maps the resulting messages or states to atomic site energies.The readout may use the final message or outputs from all previous layers.
A. Coupling of channels
Multi-ACE exposes channel coupling as a central design choice that spans existing equivariant interatomic potentials. Its parameters describe alternative routes from low-correlation, many-layer models to high-correlation, few-layer models.
- Channel interactions in the product basis affect feature-count scaling and are therefore an essential Multi-ACE design choice.
- The framework identifies a spectrum between full coupling in linear ACE and no coupling in NequIP.
- Multi-ACE characterizes models using layers, correlation order, internal and messaging spherical-harmonic orders, and feature-basis choices.
- SchNet, DimeNet, and NequIP occupy distinct parameter settings within the Multi-ACE framework.
- The formalism also accommodates Cartesian equivariant models by relating vectors to l = 1 spherical tensors.
- Existing models follow two main routes: few layers with high local correlation order, or many layers with low local correlation order.
C. Message passing as a chemically inspired sparsification
Message passing propagates information through chains of local neighbors, making it a sparse, chemically motivated alternative to a fully local ACE expansion at an enlarged cutoff. BOTNet retains selected NequIP components while enforcing body ordering through a mostly linear architecture and a final nonlinear readout.
- Message passing as sparsification: Two message-passing iterations can reach distance 2rcut, but only atoms connected through chains of closer intermediates contribute.
- Message passing as sparsification: The resulting chain-wise interaction differs from ACE, whose three-body correlations combine neighbors directly around the central atom.
- Message passing as sparsification: Under linearity, a T-layer MPNN corresponds to a sparsified one-layer ACE model with rcut,ACE = T × rcut,MPNN.
- Message passing as sparsification: A fully local ACE model at the equivalent enlarged cutoff is typically impractical because its neighborhood contains many atoms.
- BOTNet architecture: BOTNet’s features have exact correlation order t, corresponding to body-order t + 1, and its energy is expressed as a body-ordered expansion.
- BOTNet architecture: Its interaction blocks combine node features, radial features, spherical harmonics, and chemical attributes before neighborhood pooling and updates.
- BOTNet architecture: BOTNet preserves body ordering by removing intermediate nonlinearities except at the final readout, where higher-order residual terms are represented.
VI. Datasets
The computational datasets are publicly available through the BOTNet-datasets repository.
- The datasets used in the computational experiments are available at the BOTNet-datasets GitHub repository.
A. Ethanol and Methanol
The ethanol and methanol experiments use molecular-dynamics geometries to evaluate in-distribution behavior, bond-breaking extrapolation, and element-embedding effects. Related experiments examine temperature extrapolation, dihedral smoothness, and model-design consequences across benchmark settings.
- Dataset construction: The ethanol and methanol data include 500 K ab initio molecular-dynamics training geometries and independent same-distribution test configurations.
- Evaluation tasks: Removing an alcohol-group hydrogen enables a bond-breaking extrapolation test beyond the training distribution.
- Dataset construction: A mixed dataset adds 300 methanol geometries to 1000 ethanol geometries to analyze the potentials’ 2-body component.
- Evaluation tasks: The 3BPA experiments train on 300 K or mixed-temperature snapshots and test independently across temperatures.
- Evaluation tasks: Dihedral-rotation tests probe the smoothness and accuracy of the potential-energy surface governing molecular conformers.
- Evaluation tasks: Acetylacetone is deliberately evaluated with only 500 training configurations to make distinctions between models particularly challenging.
- Element embeddings: Increasing the uncoupled chemical-channel embedding size substantially changes parameter counts, while over-parameterized models often improve both in-domain and high-temperature extrapolation results.
- Element embeddings: Element embeddings also support tests of alchemical learning through dimer dissociation curves for element combinations absent from the joint training set.
Radial basis
The radial basis controls how equivariant features resolve spatial information, with element-dependent and element-agnostic choices offering different accuracy and extrapolation behavior. The section also examines body ordering and nonlinearities within the unified architecture.
- Radial basis: NequIP uses a learnable, element-agnostic radial basis conditioned on channel and symmetry indices, improving flexibility in spatial resolution.The basis uses Bessel polynomials and a cutoff function, without enforcing orthogonality.
- Radial basis: BOTNet uses separate learnable radial bases for each chemical embedding channel and neighboring chemical element.The element dependence is implemented through a weight array indexed by channels, basis functions, and product paths.
- Radial basis: Element-dependent radial bases improve training and validation accuracy, whereas element-agnostic bases perform better for extreme extrapolation such as bond breaking when correctly normalized.The comparison separates near-data accuracy from behavior far outside the training distribution.
- Non-linear Activations: Body ordering is an explicit low-dimensional structure, and preserving it in equivariant models requires linear update and readout functions or suitable finite-expansion nonlinearities.General nonlinearities can introduce infinite body order, while finite Taylor expansions such as squared norms preserve body ordering.
- Non-linear Activations: BOTNet keeps its first five message-passing layers body-ordered and uses a nonlinear final readout to represent residual higher-order contributions.This decomposes the energy into a low-body-order part and a residual term.
- Non-linear Activations: In NequIP, replacing SiLU with tanh significantly worsens results, while adding a nonlinear layer improves BOTNet beyond a strictly body-ordered model.The paper attributes tanh’s poorer optimization to vanishing gradients for large positive and negative inputs.
C. Self-Connection
Self-connections mix prior-layer information with current messages and re-inject central-atom chemical information. Their design affects message-passing accuracy, isolated-atom energy behavior, and potential-energy smoothness.
- C. Self-Connection: Self-connections mix information from the previous layer with the current layer’s output through a learnable mechanism related to residual architectures.In NequIP, this mechanism compensates for chemical information about the central atom being diluted across message-passing steps.
- C. Self-Connection: Applying a residual self-connection at the first update can introduce a learnable energy shift for isolated atoms.To preserve the correct isolated-atom limit, the first update should omit this self-connection.
- C. Self-Connection: The simplified self-connection excludes t = 0 features from the energy expression, removing the learnable shift while retaining chemical-information reinjection.A mixed BOTNet design closely matches the performance of an entirely residual architecture.
- C. Self-Connection: Self-connection choice is crucial for message-passing accuracy, with residual connections performing significantly better than having no residual architecture.The comparison includes NequIP and several BOTNet self-connection variants on the 3BPA dataset.
- D. Numerical stability: Using 32-bit floats produces a piecewise-linear, unsmooth potential-energy surface, while 64-bit precision significantly improves smoothness in NequIP and BOTNet.The observation comes from varying a bond angle in the 3BPA molecule.
IX. Normalization
The paper examines internal and data normalization as design choices affecting optimization, accuracy, and extrapolation. Internal normalization improves convergence, while physically informed data normalization preserves correct dissociation behavior but can trade accuracy for physical limits.
- Internal Normalization: Internal normalization is crucial for converging over-parametrized models trained with stochastic gradient estimation.It includes procedures applied to internal features and weights to enforce statistical properties.
- Internal Normalization: Message normalization improves performance, especially at high temperatures, reducing error by over 30% without changing model expressiveness.The difference is attributed to learning dynamics during optimization.
- Data Normalization: Scale shifting standardizes targets to zero mean and unit variance but introduces a non-physical potential-energy offset.This offset does not correspond to isolated-atom energies.
- Data Normalization: Scale shifting is suitable when dissociation is absent, such as bulk simulations, but can be problematic for reactive force fields.Physical normalization addresses the isolated-atom limit needed in dissociation settings.
- Data Normalization: Physical normalization uses isolated-atom energies and ensures the correct dissociated limit with no interaction energy.The scaling factor can be interpreted as a change of units.
- Data Normalization: Scale-shifted models achieve the best 3BPA accuracy, whereas physically normalized models are constrained to obey correct limits far from the training distribution.The two normalization schemes define substantially different learning tasks.
X. Benchmark Experiments
The benchmark experiments compare equivariant graph models with earlier approaches on in-domain accuracy, higher-temperature extrapolation, and flexible-molecule potential-energy surfaces. BOTNet and NequIP are generally highly accurate, while BOTNet is strongest in the most extreme 3BPA extrapolation test.
- Benchmark Overview: Equivariant graph neural networks are on average at least twice as accurate as kernel, linear, and feed-forward neural-network methods on organic-molecule potential-energy surfaces.BOTNet and NequIP achieve similar accuracy across a wide range of benchmarks.
- rMD17: The rMD17 experiments evaluate mean absolute energy and force errors using 1,000 training configurations for each molecule.The dataset contains five train-test splits covering ten small organic molecules.
- 3BPA Extrapolation: BOTNet and NequIP outperform linear ACE by about a factor of 2 on 3BPA at 300K and 600K.The 300K test measures in-domain accuracy, while 600K tests inputs farther from training data.
- 3BPA Extrapolation: At 1200K, BOTNet performs around 20% better than NequIP and over two times better than all other models.This test represents the most extreme extrapolation evaluated in the study.
- 3BPA Extrapolation: On the most challenging dihedral cut, linear ACE overestimates rotation barriers by about a factor of two, while NequIP and BOTNet reproduce barrier height accurately.The challenging cut has no nearby training data; BOTNet also captures the overall energy shift.
- 3BPA Extrapolation: Nonlinear models remain smooth and accurate far from the training distribution, whereas linear ACE can make larger errors despite smooth extrapolation.All three models perform similarly on the easier dihedral cuts.
C. Acetylacetone: flexibility and reactivity
The acetylacetone experiments test temperature and internal-coordinate extrapolation using a deliberately small training set. NequIP and BOTNet produce smooth potential-energy surfaces and accurately reproduce key barriers, while the paper links model design choices to accuracy and extrapolation.
- Experimental setup: A small training set makes acetylacetone a challenging test for distinguishing inference models rather than constructing the most accurate potential-energy surface.The experiments probe extrapolation in temperature and along two internal molecular coordinates.
- Dihedral extrapolation: Training samples dihedral angles below 30° but testing extends to 180°, probing substantial extrapolation in both input and energy space.The rotation barrier is about 1 eV, making this a demanding extrapolation.
- Dihedral extrapolation: All models produce a smooth dihedral potential-energy surface and reproduce its maximum near 90°, while NequIP and BOTNet accurately recover the barrier height.NequIP reproduces the potential-energy surface better than BOTNet after the maximum.
- Hydrogen-transfer reactivity: All models reproduce the hydrogen-transfer barrier shape, with BOTNet and NequIP placing its height within 2 meV.This reaction-coordinate test probes reactivity not too far from the training set.
- Design implications: The study identifies a broad Multi-ACE design space and highlights internal normalization and data normalization as important for accuracy and extrapolation.BOTNet retains NequIP’s equivariant tensor-product and learnable residual architecture while changing radial-basis, nonlinear-activation, and readout choices.
XIV. Appendix
The appendix explains how equivariant features transform and how nonlinear operations preserve equivariance. It distinguishes invariant-message nonlinearities from square-norm gating for general equivariant channels, while noting that SiLU admits infinite body order.
- Equivariant features: Spherical tensors transform through Wigner D-matrices, and symmetrisation can enforce the required equivariance.The appendix describes features using spherical coordinates and tensor indices.
- Invariant nonlinearities: Invariant messages remain unchanged under rotations, so applying any non-linearity to them preserves equivariance.The invariant channels satisfy rotational invariance under the stated group action.
- Equivariant nonlinearities: For general equivariant channels, square-norm gated nonlinearities preserve equivariance because the squared feature norm is an invariant scalar.The non-linearity acts on the squared norm rather than directly on the transforming feature.
- Body order: The Taylor expansion of SiLU is expressed using Bernoulli numbers, and SiLU admits infinite body order.The appendix explicitly notes the infinite body-order consequence.