Source-linked AI summary

Graph atomic cluster expansion for foundational machine learning interatomic potentials

Yury Lysogorskiy, Anton Bochkarev, Ralf Drautz

arXiv:2508.17936v2cond-mat.mtrl-sci

TL;DR

Accurate, efficient interatomic potentials spanning diverse materials remain difficult because conventional MLIPs have limited elemental scope and costly retraining requirements. The paper introduces GRACE foundational potentials trained on large materials datasets and evaluates their accuracy, efficiency, adaptability, and stability across tasks. GRACE models show strong broad-domain performance, while fine-tuning and distillation support specialized or simpler models without losing general knowledge.

  • Problem

    Conventional MLIPs have limited elemental scope, making expansion to new elements require thousands of reference calculations and full-model retraining.

  • Method

    The study develops foundational GRACE potentials using a complete graph-based atomic-cluster basis, large materials datasets, varied architectures, fine-tuning, and knowledge distillation.

  • Results

    GRACE models establish a Pareto front across accuracy and efficiency benchmarks, with strong thermal-conductivity, long-time molecular-dynamics, and specialized-task performance.

  • Takeaways & Limitations

    GRACE provides a robust and adaptable foundation for high-fidelity atomistic simulations across broad chemical and structural domains.

Abstract

from arXiv · show

Foundational machine learning interatomic potentials that can accurately and efficiently model a vast range of materials are critical for accelerating atomistic discovery. We introduce universal potentials based on the graph atomic cluster expansion (GRACE) framework, trained on several of the largest available materials datasets. Through comprehensive benchmarks, we demonstrate that the GRACE models establish a new Pareto front for accuracy versus efficiency among foundational interatomic potentials. We further showcase their exceptional versatility by adapting them to specialized tasks and simpler architectures via fine-tuning and knowledge distillation, achieving high accuracy while preventing catastrophic forgetting. This work establishes GRACE as a robust and adaptable foundation for the next generation of atomistic modeling, enabling high-fidelity simulations across the periodic table.

I. INTRODUCTION

Conventional MLIPs face limited elemental scope and costly retraining, motivating universal models that cover the periodic table efficiently. This work develops GRACE potentials using large materials datasets and multiple model complexities.

  • Motivation: Adding a new element to conventional MLIPs requires thousands of reference calculations and retraining the entire model.This creates a significant computational bottleneck.
  • Motivation: Universal potentials must represent vast chemical interactions without brute-force enumeration, which would require about 10^8 four-body interactions.Foundational MLIPs instead use low-dimensional chemical embeddings to capture periodic-table chemistry.
  • GRACE framework: GRACE extends ACE to tree-graphs, providing a complete basis for atomic interactions and sparse chemical embeddings evaluated recursively.The recursive basis evaluation can be understood as message passing.
  • Validation challenge: Validation of universal MLIPs is inherently limited because combinatorially many simulations across the periodic table cannot be tested.Evaluations therefore focus on particularly relevant or widely adopted tests.
  • Models and data: The study parameterizes GRACE models on OMat24, Alexandria, and MPTraj, including one- and two-layer architectures with small, medium, and large complexities.OMat24 contains 110 million DFT calculations and spans 89 elements; fine-tuning with MPTraj and sAlex produces OAM models designed to avoid WBM test leakage.

B. Validation

The validation suite tests foundational MLIPs on crystal discovery, thermal conductivity, and force-dependent properties. GRACE models consistently occupy the Pareto front, with GRACE-2L-OAM-L achieving the lowest reported thermal-conductivity error.

  • Scope: The validation suite covers formation energies and stability in MatBench Discovery, thermal conductivity through κSRME, and additional non-equilibrium and defective configurations.MatBench Discovery optimizes 257,000 WBM candidate crystals spanning diverse compositions.
  • Materials discovery: GRACE models consistently occupy the Pareto front for stable-structure identification, balancing performance against computational speed.Figure 1 compares F1 score and computational time per atom, with higher F1 indicating better performance.
  • Thermal conductivity: κSRME evaluates thermal-conductivity prediction across 103 binary structures using forces processed with phono3py.The metric probes force-dependent properties involving phonons and anharmonic thermal conductivity.
  • Thermal conductivity: κSRME = 0.168 was achieved by GRACE-2L-OAM-L, the lowest error among the compared foundational MLIPs.The one- and two-layer GRACE models formed the Pareto front for thermal-conductivity prediction.

3. Elastic properties

Elastic and defect-property tests assess GRACE and other foundational MLIPs using relative and absolute error metrics. GRACE-2L-OAM-L has the lowest elastic-constant CSRME, while defect errors vary by property and element.

  • Elastic properties: Elastic constants are grouped into longitudinal, Poisson’s ratio-related, and shear subgroups, with SRME and MAE reported against Materials Project references.The subgroups differ in magnitude, motivating use of both relative and absolute errors.
  • Elastic properties: GRACE-2L-OAM-L demonstrated the lowest CSRME among all tested models.Longitudinal constants generally have lower relative but higher absolute errors, whereas the other groups show the opposite pattern.
  • Grain boundaries: Grain-boundary formation-energy accuracy is evaluated with γGB-SRME and mean absolute error ∆γGB for pure metals.These metrics quantify relative and absolute errors for bulk structural defects.
  • Grain boundaries: γGB-SRME generally ranges from 0.275 to 0.4, while ∆γGB remains below 5 meV/˚A2 for most models.K, Rb, and Cs show consistently larger relative errors, suggesting that a 6 ˚A cutoff may be insufficient for these alkali elements.
  • Surfaces: Surface-energy errors span γsurf-SRME values from 0.168 to 0.279 and typical absolute errors of 8 to 14 meV/˚A2.The reported range covers the tested foundational MLIPs for pure-element surfaces.

6. Point defects

GRACE models are evaluated on unary point-defect formation energies using symmetric relative mean errors, with both defect types generally showing moderate errors and corresponding absolute-error ranges.

  • The evaluation uses ESIA-SRME and Evac-SRME because formation energies vary in scale across defect types and metals.
  • SIA and vacancy formation-energy SRME values generally fall between 0.1 and 0.3, with a few outliers.
  • SIA formation-energy MAE typically ranges from about 0.2 to 0.4 eV, while vacancy formation-energy MAE ranges from 0.1 to 0.2 eV.

D. Computational performance

GRACE models provide scalable computational performance across diverse materials and system sizes, with simpler one-layer models offering greater speed and memory capacity than two-layer models.

  • The LAMMPS evaluation used carbon diamond, liquid water, aluminum FCC, and molten FLiBe systems with incrementally increased sizes until out-of-memory errors.
  • 27–120 µs/atom per step for two-layer models and 10–28 µs/atom per step for single-layer models were measured across systems.For the most intensive GRACE-2L-L model, up to 20,000 carbon atoms ran at approximately 124 µs per atom per step.
  • 20,000–55,000 atoms fit on one A100-80GB GPU for two-layer models, versus 78,000–215,000 atoms for one-layer models.One-layer locality permits domain decomposition in LAMMPS for very large systems.
  • GRACE models also showed strong performance on commodity RTX3060 GPUs despite consistently using FP64 precision.

2. Hydrogen combustion

Fine-tuning GRACE-2L-OAM on hydrogen-combustion data exposes a trade-off between downstream accuracy and retention of the original task. Frozen-weight strategies mitigate catastrophic forgetting under the dataset shift.

  • The H2COMB dataset contains IRC, ab initio MD, and normal-mode calculations covering 19 hydrogen-combustion reaction channels.
  • Naive fine-tuning reaches 37 meV/˚A force MAE on H2COMB but increases original sAlex errors to 339 ×10^3 meV/˚A for H/O and 6.5 ×10^3 meV/˚A for other elements.
  • Frozen-weight strategies limit sAlex force-MAE increases for elements excluding H and O from 25 meV/˚A to 31 and 35 meV/˚A.For structures containing neither H nor O, predictions remain unchanged from baseline.
  • Freezing selected model components mitigates forgetting even for updated elements despite systematic differences between the foundational and H2COMB reference data.

F. Model distillation

GRACE distillation trades a small amount of teacher accuracy for substantially lower computational cost, while extended distillation restores general chemical performance more effectively than standard pathways.

  • The primary task measures HEA25S force-component MAE, while the secondary task measures formation-energy MAE for non-magnetic Materials Project binaries.
  • The original foundation model has large primary-task error from the PBE-to-PBEsol mismatch but retains 14 meV/atom secondary-task error.
  • Naive distilled, distilled/finetuned, and bespoke models have similar error metrics, with the bespoke model slightly better on the primary task.
  • Extended Distilled substantially improves the secondary task toward teacher performance while keeping primary-task metrics only slightly worse than the bespoke model.
  • 0.56 ms/atom/core for GRACE-FS versus 34.48 ms/atom/core for GRACE-2L yields a speedup of nearly 70×.Student models show slight accuracy degradation because of their simpler architecture.

III. DISCUSSION

GRACE foundational potentials combine a formally complete atomic basis with broad DFT training and extensive validation. They support accurate, stable, and adaptable modeling across diverse structures and specialized tasks.

  • III. DISCUSSION: GRACE potentials use a formally complete basis with exact physical symmetries, conservative forces, and FP64 evaluation.They guarantee invariance under rotation, inversion, translation, and permutation of identical species, with forces computed as energy gradients.
  • III. DISCUSSION: The models were trained on a massive DFT dataset spanning broad atomic configurations and chemistries.
  • III. DISCUSSION: GRACE models largely define the MatBench Discovery Pareto front across thermodynamic stability, formation-energy MAE, and computational speed.Two-layer models also show leading thermal-conductivity performance.
  • III. DISCUSSION: Validation on grain boundaries, surfaces, and point defects shows that GRACE describes non-equilibrium and open structures effectively.Outliers for K, Rb, and Cs are attributed to a cutoff radius that is too small for these large elements.
  • III. DISCUSSION: Long-time FLiBe molecular dynamics simulations maintain stability while reproducing dynamic properties including radial distribution functions and diffusion coefficients.
  • III. DISCUSSION: Fine-tuning and distillation adapt foundational GRACE models to specialized tasks and simpler architectures, including low-data settings and catastrophic-forgetting mitigation.Freezing selected layers mitigates loss of general knowledge, while distillation improves accuracy over a wider configurational space compared with training from scratch.

B. Validation

The validation program spans computational performance, elastic properties, grain boundaries, surfaces, and point defects. These evaluations use standardized datasets and common subsets where necessary to enable comparisons.

  • B. Validation: Computational performance was measured from averaged wall time for ten molecular-dynamics steps on a 1024-atom tungsten BCC supercell using a single NVIDIA A100 GPU.Initial runs were excluded for JIT compilation and caching effects, and each simulation was repeated independently ten times.
  • B. Validation: Elastic constants were compared with Materials Project references using a common subset of 7,962 structures successfully evaluated by all potentials.The original reference set contained 10,073 elastic-matrix calculations, but some structures failed because of relaxation or convergence issues.
  • B. Validation: Grain-boundary validation used Crystalium structures and standardized BFGS and FIRE relaxation procedures with force-based stopping criteria.The initial dataset contained 327 grain boundaries across 58 elements.
  • B. Validation: Surface validation used Crystalium structures with the same relaxation settings, selecting 716 structures whose surface planes were orthogonal to the z-direction.These were selected from an initial set of 1,124 surface structures.
  • B. Validation: Self-interstitial and vacancy formation energies were taken from reference studies covering 13 BCC and 13 FCC metals, all computed with the PBE functional.

A. Models architecture and parameterization

GRACE represents atomic environments with chemical embeddings and recursively evaluated ACE basis functions, using one- and two-layer architectures with different complexity levels. Models are parameterized on combined materials datasets and evaluated through fine-tuning and distillation pathways.

  • A. Models architecture and parameterization: GRACE uses chemical species, bond vectors within a 6 Å cutoff, Chebyshev radial functions, and spherical harmonics up to lmax = 44.
  • A. Models architecture and parameterization: Recursive sparse ACE evaluation constructs many-body basis functions through equivariant sums and Clebsch–Gordan coupling operations.Two-layer models linearly combine equivariant representations that extend the effective interaction range to 2rcut before a second recursive evaluation.
  • A. Models architecture and parameterization: The models were parameterized using OMat24 together with sAlex and MPTraj, using raw VASP energies, forces, and stresses as targets.
  • A. Models architecture and parameterization: Training minimizes weighted energy, force, and stress losses, first on OMat24 and then through fine-tuning on MPTraj and sAlex.The loss uses a Huber form, with stage-specific weights for energy, forces, and stresses.
  • A. Models architecture and parameterization: Fine-tuning and distillation pathways were evaluated using the configurations summarized in Table S2 and Figure S2.
  • A. Models architecture and parameterization: Including atomic site energies as additional distillation targets produced slightly worse metrics than the otherwise identical parameterization without them.

C. Molecular dynamics stability and energy conservation

GRACE shows stable long-time molecular-dynamics behavior in molten FLiBe, with total-energy fluctuations monitored over a 1 ns NVE trajectory at 973 K. The associated figures also organize accuracy–efficiency trade-offs across architectures and training strategies.

  • C. Molecular dynamics stability and energy conservation: 5 · 10^-9 eV/atom/ns is the reported linear total-energy drift during a 1 ns FLiBe NVE simulation at 973 K.The energy-fluctuation histogram is described as symmetric and Gaussian-like, indicating a stable equilibrium distribution.
  • C. Molecular dynamics stability and energy conservation: Figure S1 presents the GRACE overall scheme and recursive ACE basis evaluation in top and bottom panels, respectively.
  • C. Molecular dynamics stability and energy conservation: Figure S2 plots primary-task force MAE against secondary-task formation-energy MAE, with lower values indicating better accuracy on both axes.
  • C. Molecular dynamics stability and energy conservation: Computational cost is shown as wall-clock time per atom step against primary and secondary-task accuracy, with a dashed Pareto front marking optimal speed–accuracy trade-offs.Symbols encode architecture and model size, colors encode training methodology, and empty versus filled symbols distinguish 1L- and 2L-teacher distilled models.

D. Detailed Validation Benchmarks

This section surveys detailed validation benchmarks spanning formation-energy accuracy, computational performance, elastic tensors, and structural defects. It also reports elementwise error decompositions for grain boundaries, surfaces, and unary point defects.

  • Formation energy and efficiency: Validation covers MatBench Discovery formation-energy MAE versus computational time per atom, with Pareto-optimal models linked and GRACE models marked in red.Performance is estimated using ASE and LAMMPS.
  • Elastic tensors: Elastic-tensor performance is evaluated using SRME and MAE across three subgroups referenced to Materials Project.
  • Structural defects: Unary grain-boundary and surface formation energies are assessed with SRME and mean absolute error metrics, with GRACE models highlighted.Elementwise decompositions are provided for both structural-defect classes.
  • Elementwise analysis: The benchmark suite includes elementwise SRME decompositions for grain boundaries, surfaces, self-interstitials, and vacancies across elements and machine-learning interatomic potentials.
  • Structural defects: Unary point-defect accuracy is reported separately for self-interstitials and vacancies using SRME-based error metrics.Element-decomposed SRME results are given for both self-interstitials and vacancies.
Loading 2508.17936v2…