Source-linked AI summary
Evaluation of the MACE Force Field Architecture: from Medicinal Chemistry to Materials Science
David Peter Kovacs, Ilyes Batatia, Eszter Sara Arany, Gabor Csanyi
TL;DR
Machine learning force fields need accurate, transferable models across diverse chemical and physical systems, including low-data and large-molecule settings. This paper benchmarks MACE on published datasets using a local, equivariant many-body architecture and a general training strategy. MACE generally outperforms alternatives, remains effective with limited data, and describes large molecular assemblies despite its strictly local construction.
Problem
The paper addresses the need to evaluate whether one force-field architecture applies accurately across diverse chemical and physical systems, including low-data and large-molecule tasks.
Method
The authors train MACE on published benchmark datasets and use its local equivariant many-body architecture with a general loss-scheduling training strategy.
Results
MACE generally outperforms alternatives across tested systems and tasks, including large molecules, small-molecule torsions, condensed phases, and low-data molecular fitting.
Takeaways & Limitations
MACE provides a broadly applicable and data-efficient force-field approach, with strictly local models sufficient for large molecules and weakly interacting assemblies in the tested settings.
Abstract
from arXiv · showhide
The MACE architecture represents the state of the art in the field of machine learning force fields for a variety of in-domain, extrapolation and low-data regime tasks. In this paper, we further evaluate MACE by fitting models for published benchmark datasets. We show that MACE generally outperforms alternatives for a wide range of systems from amorphous carbon, universal materials modelling, and general small molecule organic chemistry to large molecules and liquid water. We demonstrate the capabilities of the model on tasks ranging from constrained geometry optimisation to molecular dynamics simulations and find excellent performance across all tested domains. We show that MACE is very data efficient, and can reproduce experimental molecular vibrational spectra when trained on as few as 50 randomly selected reference configurations. We further demonstrate that the strictly local atom-centered model is sufficient for such tasks even in the case of large molecules and weakly interacting molecular assemblies.
I. Introduction
The paper evaluates MACE across chemical and physical systems using published benchmark datasets, while introducing its local, equivariant, many-body architecture and general training strategy.
- I. Introduction: MACE is benchmarked on a wide variety of tasks to assess out-of-the-box applicability across the chemical and physical sciences.The evaluation uses previously published datasets and spans molecular, materials, condensed-phase, and benchmark-target prediction tasks.
- A. Many-body equivariant message passing: MACE maps atomic positions and chemical elements to total potential energy by decomposing it into atom-centered site energies.Each site energy depends on symmetric features describing the atom’s chemical environment.
- A. Many-body equivariant message passing: Local neighborhoods include atoms within a predefined cutoff, and spherical-harmonic node features provide rotational equivariance through message passing.The model initializes element embeddings, combines neighbor displacement information with radial and angular bases, and iterates through layers.
- A. Many-body equivariant message passing: MACE forms many-body features by tensor products, generalized Clebsch-Gordan contractions, and learnable messages exchanged between neighboring atoms.The maximum body-order is controlled by the tensor-product order, while multiple layers expand the effective receptive field to approximately S × rcut.
- A. Many-body equivariant message passing: Forces are obtained as analytical derivatives of the total potential energy using autodifferentiation tools.The model’s readout uses rotationally invariant node features, with linear readouts preserved except where the final-layer design differs.
B. The body-order of MACE models
MACE represents local environments with body-ordered descriptors whose layered construction efficiently generates high-body-order features.
- B. The body-order of MACE models: A two-layer MACE model reaches 13-body features from layers with body-order 4, supporting efficient high-body-order representation.The first layer produces 4-body functions; the second adds one-particle information and then forms tensor products three times.
- B. The body-order of MACE models: The descriptors can linearly span symmetric functions over chain-like clusters up to the model’s maximum body-order in the complete-basis limit.The relevant clusters extend through graph hops of size rcut.
C. Loss scheduler
MACE training combines energy and force errors with scheduled weights so that force accuracy is retained while absolute energy errors are reduced.
- C. Loss scheduler: The loss combines weighted mean squared errors for total energies, force components, and available virials or stresses.Energy and force weights are controlled by λE and λF.
- C. Loss scheduler: High force weighting improves force accuracy but does not necessarily ensure accurate energies, especially for heterogeneous training sets.Heterogeneous sets contain widely separated systems in atomic configuration space, such as different molecules or solid phases.
- C. Loss scheduler: For about 60% of training, λF > λE; later, λE > λF and the learning rate decreases by a factor of 10.The schedule is designed to reduce absolute energy errors while preserving high force accuracy.
III. Locality of large molecular systems
The paper tests MACE’s locality assumption on large molecules and molecular assemblies with hundreds of atoms and complex intermolecular interactions.
- III. Locality of large molecular systems: MACE is compared with global sGDML and mixed short-range/long-range VisNet-LSRM models on these systems.The comparison examines how local message passing performs against models using global or fragment-level descriptions.
- III. Locality of large molecular systems: MD22 contains large molecules and molecular assemblies with hundreds of atoms and complex intermolecular interactions.Its configurations were sampled using elevated-temperature ab initio molecular dynamics simulations.
- III. Locality of large molecular systems: Table I reports energy and force MAEs in meV/atom and meV/Å, respectively, alongside approximate system diameters.Rows compare global, long-range, and local model classes on the MD22 dataset.
A. Effect of locality on energy and force errors
MACE’s locality–accuracy trade-off depends on the receptive field: a 2×5 Å model improves sGDML errors by up to tenfold, while 3 Å layers can miss intermolecular interactions. Longer-range local models nevertheless reproduce qualitatively correct dynamics and vibrational spectra for weakly interacting assemblies.
- Energy and force errors: Up to a factor of 10 improvement over sGDML errors is achieved by MACE with a 2×5 Å cutoff.Even the 2×3 Å and 1×6 Å models outperform sGDML for all tested systems.
- Comparison with mixed-range models: The strictly local MACE model has significantly lower force errors than VisNet-LSRM, although the long-range model usually has lower energy errors.Energy errors for both models are generally 0.1 meV/atom or lower.
- Energy and force errors: 5-6 Å receptive fields appear sufficient to capture typical intermolecular interactions, whereas 3 Å layers produce larger energy errors in nucleic-acid and Bucky-ball catcher systems.The shorter-range model cannot describe interactions whose separations typically exceed 3 Å.
- Bucky-ball catcher dynamics: The 2×3 Å Bucky-ball catcher model dissociates within 10 ps despite 0.5 meV/atom energy and 13 meV/Å force accuracy.The two longer-range MACE models provide qualitatively correct dynamics.
- Bucky-ball catcher dynamics: A local 1×6 Å MACE model shows excellent agreement with the sGDML molecular vibrational spectrum, including low frequencies.This comparison demonstrates that local models can simulate the dynamics of the Bucky-ball catcher when the receptive field is sufficiently long.
A. Biaryl torsion benchmark
The biaryl torsion benchmark tests MACE on 88 drug-like molecules where accurate torsional barriers are difficult for classical force fields. Transfer-learned MACE models substantially reduce barrier-height errors relative to ANI-1ccx and reproduce challenging torsional profiles smoothly.
- Benchmark scope: 88 small drug-like molecules with biaryl dihedral torsional profiles comprise the challenging benchmark dataset.Accurate torsional barriers are relevant to small-molecule drug discovery because classical empirical force fields often describe them insufficiently.
- Benchmark results: 0.36 kcal/mol is the mean absolute barrier height error of MACE 192-2 versus 0.78 kcal/mol for ANI-1ccx.ANI-1ccx was identified as the best model to date in the benchmark comparison.
- Model-size comparison: 0.56 and 0.83 kcal/mol are the average barrier-height errors of the medium and small MACE models, respectively.Both models perform better than or comparably to ANI-1ccx.
- Torsional scans: All three shown challenging torsional scans have smooth MACE surfaces with the correct positions of minima and maxima.The reported behavior is consistent with the remaining test molecules available in the Supplementary Information.
VI. Amorphous carbon
MACE is evaluated on diverse carbon phases and broad condensed-material datasets, achieving strong accuracy across chemically varied systems. The training strategy is adapted to heterogeneous data through robust losses and generalized radial features.
- Amorphous carbon: MACE 256-2 significantly improves test-set errors over the state-of-the-art ACE potential across all carbon phases.The carbon dataset spans crystalline, amorphous, and clustered structures with varied bonding environments.
- HME21: 37 chemical elements make HME21 a demanding benchmark for simultaneous force-field modelling across diverse chemical space.The dataset contains both disordered and regular crystals.
- HME21: MACE outperforms NequIP and TeaNet by more than 30% in both energy and force metrics on HME21.The reported low errors indicate strong accuracy across the dataset’s broad elemental range.
- M3GNet: The M3GNet dataset targets a universal condensed-matter potential spanning all 89 elements from Hydrogen to Thorium.Its materials are extracted from the Materials Project and cover an extensive range of systems.
- M3GNet: For the diverse M3GNet dataset, MACE training uses Huber loss and an element- and environment-dependent radial basis to tolerate outliers.The Huber loss switches between L1 and L2 behaviour.
IX. Liquid Water
MACE performs strongly on liquid-water force-field fitting and reproduces key thermodynamic and dynamical behaviour in molecular simulations. It also achieves state-of-the-art performance on several QM9 tasks, with accuracy improved by exploiting zero-force information.
- Liquid-water fitting: 1593 liquid-water configurations containing 64 molecules each form the training dataset, computed with CP2K at the revPBE0-D3 DFT level.The reference method is reported to describe water structure and dynamics reasonably well across pressures and temperatures.
- Liquid-water fitting: A relatively small invariant MACE model with overall body-order 13 achieves lower water errors than other best models, while the larger MACE model only slightly improves force errors.Table V compares energy and force errors against BPNN, REANN, and NequIP trained on the same dataset with different splits.
- Liquid-water dynamics: 2.14 ± 0.14 × 10^-9 m2/s is the MACE water diffusion coefficient, in reasonably good agreement with an ab initio estimate.The value comes from three independent 200 ps NVT simulations at equilibrium density 0.91 g / cm3.
- QM9 benchmark: MACE achieves state-of-the-art results on 3 of 4 energy-related QM9 tasks and improves the state of the art on one additional non-energy task.The benchmark contains 12 tasks overall.
- QM9 learning curves: Including zero-force information from equilibrium QM9 geometries increases the accuracy of both small and large MACE models without new quantum-mechanical calculations.The effect is shown in the potential-energy learning-curve analysis.
- QM9 learning curves: Higher body-order features distinguish MACE, Allegro, and Wigner kernels from several models with comparable but higher QM9 errors.The comparison suggests high body-order is important among the most successful atomistic machine-learning models.
XI. Conclusion
Across molecular, condensed-phase, and quantum-chemistry benchmarks, MACE delivers accurate and transferable force fields with limited data and minimal task-specific modification. The results support strictly local models for large systems and broad applicability across chemical and physical domains.
- Locality and large molecules: A two-layer MACE model with a 5 Å cutoff per layer improves accuracy over a state-of-the-art global model by up to a factor of 10.The paper also reports accurate molecular vibrational spectra from the local model.
- Transferability and data efficiency: A MACE model trained on only 50 reference calculations reproduces the coupled-cluster-level vibrational spectrum of ethanol without iterative training.The study also demonstrates transfer learning to coupled-cluster theory.
- Cross-domain conclusions: MACE achieves excellent accuracy and transferability on large, chemically diverse materials datasets and improves the state of the art on many QM9 properties.The conclusion covers carbon, liquid water, universal materials modelling, and molecular benchmarks.
- Cross-domain conclusions: The authors report that MACE handles varied systems with little to no modification to training and hyperparameters.They identify this as a route toward high-accuracy force fields requiring limited user input or expertise.
Appendix
The appendix reports that a two-phase loss schedule improves energy accuracy while preserving force accuracy, alongside alternative loss weighting and radial-feature choices.
- Two phases of learning: Changing loss weights dramatically improves energy errors without significantly affecting force accuracy.The result is shown for the medium 96-1 MACE model.
- Two phases of learning: The two-phase loss schedule decreases energy and force validation errors compared with using only the second loss phase.
- Loss function: The alternative loss averages energy per atom and forces per atom squared, producing force-to-energy weights up to 1,000:1 instead of the typical 10:1 ratio.All models used this alternative loss except where otherwise stated.
- Radial representation: For datasets containing many elements, the radial representation can be conditioned on sender and receiver atom features rather than remaining element-agnostic.The radial MLP receives the scalar parts of both atoms’ preceding node features.
D. Computational details
The computational details specify shared MACE architecture and optimization settings, with dataset-specific configurations for molecular, condensed-phase, and materials benchmarks.
- Training procedure: Training took 1–5 hours for small molecules and up to 3–4 days for the largest ANI-1x and QM9 models, stopping when validation loss ceased improving.
- Common settings: All common models used two MACE layers, lmax = 3, eight radial Bessel features, correlation order N = 3, AMSGrad, and an on-plateau learning-rate scheduler.The models also used exponential moving averaging with weight 0.99 and the loss from Equation (B.1).
- Dataset-specific models: MD22 models used 256 uncoupled channels, a 95%/5% train-validation split, energy mean shifting, and a two-phase switch from λE = 10, λF = 1000 to λE = 1000, λF = 10.
- Dataset-specific models: COMP6 models used a 5 Å cutoff, 64–192 channels, Lmax = 0–2, and 196,944–2,229,456 parameters across increasing model sizes.
- Dataset-specific models: The liquid-water models used a 6 Å cutoff, while the larger model used 192 channels and Lmax = 2.The small model used 64 uncoupled channels and invariant messages.
- Dataset-specific models: For intensive QM9 properties, nonlinear pooling with sum, mean, and standard-deviation features plus attention improved error by up to twofold over simple linear extensive pooling.
E. Effect of locality on molecular dynamics simulations
The locality tests indicate that short-range MACE models can reproduce vibrational spectra for larger molecules, while broader molecular tests show mostly low energy errors with systematic shifts in two larger peptides.
- Tetrapeptide: Short-range and longer-range MACE models agree remarkably well on the tetrapeptide vibrational spectrum.This supports using short-range models for accurate vibrational spectra of larger systems.
- Molecular energy correlations: 12 of 14 COMP6 molecules have very low errors near the DFT values, while two larger peptides show a systematic shift from ground truth.
- Low-data vibrational spectra: Models trained on only 50 QM calculations reproduce accurate vibrational spectra for salicylic acid and paracetamol.These molecules are described as more challenging test cases than ethanol.