Source-linked AI summary

PET-MAD, a lightweight universal interatomic potential for advanced materials modeling

Arslan Mazitov, Filippo Bigi, Matthias Kellner, Paolo Pegolo, Davide Tisi, Guillaume Fraux, Sergey Pozdnyakov, Philip Loche, Michele Ceriotti

arXiv:2503.14118v2cond-mat.mtrl-scics.LGphysics.chem-ph

TL;DR

Universal MLIPs offer efficient alternatives to first-principles simulations but can be limited by incomplete coverage of the vast structural and compositional space. PET-MAD combines a lightweight transformer-based potential with a diverse, internally consistent dataset and demonstrates competitive accuracy across benchmarks and advanced simulations, while supporting uncertainty quantification and targeted fine-tuning.

  • Problem

    Universal force fields face epistemic uncertainty because the vast structural and compositional space is only sparsely covered by training data.

  • Method

    PET-MAD combines an unconstrained transformer-based graph neural network with the diverse MAD dataset, consistent reference energetics, LLPR uncertainty quantification, and LoRA fine-tuning.

  • Results

    PET-MAD provides semi-quantitative accuracy across inorganic and organic benchmarks and six diverse applications despite using a much smaller dataset than current state-of-the-art models.

  • Takeaways & Limitations

    PET-MAD offers a compact framework for accurate, fast, uncertainty-aware atomistic simulations and can be further improved for target systems with limited additional training data.

Abstract

from arXiv · show

Machine-learning interatomic potentials (MLIPs) have greatly extended the reach of atomic-scale simulations, offering the accuracy of first-principles calculations at a fraction of the cost. Leveraging large quantum mechanical databases and expressive architectures, recent ''universal'' models deliver qualitative accuracy across the periodic table but are often biased toward low-energy configurations. We introduce PET-MAD, a generally applicable MLIP trained on a dataset combining stable inorganic and organic solids, systematically modified to enhance atomic diversity. Using a moderate but highly-consistent level of electronic-structure theory, we assess PET-MAD's accuracy on established benchmarks and advanced simulations of six materials. Despite the small training set and lightweight architecture, PET-MAD is competitive with state-of-the-art MLIPs for inorganic solids, while also being reliable for molecules, organic materials, and surfaces. It is stable and fast, enabling the near-quantitative study of thermal and quantum mechanical fluctuations, functional properties, and phase transitions out of the box. It can be efficiently fine-tuned to deliver full quantum mechanical accuracy with a minimal number of targeted calculations.

I. INTRODUCTION

Universal MLIPs reduce the cost of first-principles simulations but often depend on system-specific training and datasets biased toward stable materials. PET-MAD addresses this through a chemically and structurally diverse, internally consistent dataset spanning organic and inorganic systems.

  • Motivation: Universal MLIPs replace costly first-principles energy and force evaluations, but their accuracy depends on how closely applications match their reference datasets.System-specific potentials require substantial ab initio calculations and fitting effort for each new system.
  • Approach: PET-MAD combines the PET transformer-based architecture with the Massive Atomistic Diversity dataset and internally consistent reference energetics.The dataset is designed to include organic and inorganic systems across dimensionalities while treating structural and chemical motifs consistently.
  • Design principle: The dataset construction emphasizes consistent electronic-structure calculations so that one coherent structure-energy mapping applies across material classes.This consistency comes at the cost of neglecting spin polarization, electron correlation, and dispersion effects important for some materials.
  • Dataset: The MAD dataset contains 95595 structures spanning 85 elements and eight subsets, including bulk crystals, distorted and randomized structures, surfaces, clusters, two-dimensional crystals, molecular crystals, and molecular fragments.These subsets systematically broaden structural, chemical, and dimensional diversity.

B. Model architecture and training

PET-MAD uses a compact transformer-based graph neural network and a diverse, consistently referenced training strategy to balance accuracy, speed, and broad transferability. Benchmarks show competitive or leading performance across inorganic, molecular, and diverse out-of-distribution configurations, with data-efficiency and computational advantages over larger models.

  • Model architecture and training: PET-MAD uses an unconstrained transformer-based graph neural network with approximately 3.3 million parameters, trained on merged 80%/10%/10% training, validation, and test splits.The architecture was selected through hyperparameter search, and LoRA fine-tuning adapts it to specific chemical systems.
  • Benchmarking: Consistent DFT references are essential: PET-MAD’s inconsistent-reference WBM error is three times larger, while the discrepancy is attributed entirely to reference-data inconsistency.The benchmark subset was recomputed using MAD-compatible settings for fair model comparisons.
  • Data efficiency: PET-MAD reaches the accuracy of SevenNet and GNoME using 1-3 orders of magnitude less data on the energy-above-hull benchmark.The comparison uses compatible DFT references and evaluates data efficiency along the Pareto frontier.
  • Benchmarking: PET-MAD achieves high energy and force accuracy, outperforming MACE-MP-0-L in most cases and closely competing with MatterSim-5M and SevenNet-l3i5 on inorganic datasets.It outperforms the other models on SPICE and MD22 and matches Orb-v2 on the extrapolative OC2020 S2EF dataset.
  • Benchmarking: On the consistently recomputed MAD benchmark, PET-MAD outperforms MACE-MP-0-L, MatterSim-5M, and SevenNet-l3i5 in almost all cases, especially for molecular systems and unusual distorted configurations.Other models’ errors become up to 50 times larger on highly diverse distorted subsets, while PET-MAD dramatically outperforms them on molecular systems.
  • Computational performance: PET-MAD is competitive in speed and accuracy with larger models trained on much larger datasets, supporting longer and larger simulations.The authors attribute this combination to the unconstrained architecture and problem-agnostic training-set construction.

D. Ionic transport in lithium thiophosphate

The Li3PS4 case study tests PET-MAD on temperature-dependent ionic conductivity across three phases. PET-MAD agrees closely with bespoke and fine-tuned models and captures the conductivity transition associated with PS4 tetrahedra rotation.

  • System and observable: The Li3PS4 benchmark evaluates ionic conductivity σ across the α, β, and γ phases over a broad temperature range.Lithium thiophosphates are studied as solid-state battery electrolytes.
  • Model comparison: PET-MAD is compared with a bespoke PET model trained from scratch and a LoRA-fine-tuned model trained on the same system-specific dataset.The validation MAE is 4.9 meV/atom|63.9 meV/Å for PET-MAD, versus 1.2 meV/atom|35.6 meV/Å for the bespoke model and 1.3 meV/atom|36.0 meV/Å for the fine-tuned model.
  • Results: PET-MAD shows excellent agreement with the bespoke and fine-tuned models for all three phases despite its larger validation error.The figure compares PET-MAD, the fine-tuned model, and the bespoke model through the temperature dependence of σ.
  • Results: PET-MAD slightly overestimates the γ-phase temperature for transition to a high-conductivity phase but otherwise captures the conductivity change associated with PS4 tetrahedra rotation.The results agree quantitatively with the earlier study despite different DFT parameters and MLIP classes.

E. Melting point of GaAs

PET-MAD estimates GaAs melting through temperature-dependent chemical-potential differences, with ensemble reweighting used to propagate fit uncertainty. Its melting-point prediction is less accurate than bespoke and LoRA-finetuned models, while the discrepancy is attributed partly to electronic-structure reference limitations.

  • Method: The interface-pinning calculation determines the melting point where the liquid–solid chemical-potential difference becomes zero.Ensemble predictions and reweighting propagate epistemic uncertainty into the chemical-potential curves and melting-point estimate.
  • Results: The PET-MAD ensemble shows a large spread in chemical-potential predictions relative to the bespoke and LoRA-finetuned models.The spread is presented as reflecting the discrepancy between the models.
  • Results: PET-MAD predicts a GaAs melting point of 1111 ± 72 K, versus 1169 ± 3 K for PET-MAD-LoRA and 1169 ± 4 K for PET-Bespoke.The corresponding energy|force MAEs are 14.4|74.1 meV/atom|meV/Å for PET-MAD, 0.7|29.0 for PET-Bespoke, and 1.3|45.3 for PET-MAD-LoRA.
  • Limitation: The melting-point discrepancy is consistent with prior work and can reflect DFT-functional dependence, which can shift computed melting points by hundreds of kelvin.This limits direct interpretation of the model discrepancy as solely an ML-fitting error.
  • Related case study: For the CoCrFeMnNi alloy comparison, PET-MAD and PET-MAD-LoRA produce nearly identical nickel-enriched surface segregation patterns, while PET-Bespoke differs quantitatively.The bespoke model’s inconsistency is attributed to overfitting after recomputing only about 2000 of 30,000 HEA25S structures.

G. Quantum nuclear effects in liquid water

PIMD tests whether PET-MAD can reproduce quantum nuclear effects in liquid water and related fluctuation-sensitive observables. PET-MAD agrees closely with bespoke models, while retaining the underlying electronic-structure reference’s tendency to overstructure water and misrepresent its heat capacity.

  • Method: PIMD of 128 water molecules evaluates O-O and O-H pair correlations and the constant-volume heat capacity at 298 K to probe structural and quantum fluctuations.The heat-capacity calculation uses path-integral estimators that are difficult to converge, and the study compares PET-MAD with finetuned, bespoke, and empirical forcefield models.
  • Results: PET-MAD agrees excellently with bespoke models in both classical and path-integral simulations, including converged pair-correlation functions.The agreement covers the radial distribution functions and heat-capacity comparison shown for varying bead counts.
  • Results: PET-MAD inherits the reference functional’s overestimated water melting point, producing an overstructured undercooled liquid with excessive proton delocalization and ice-like heat capacity.These limitations prevent the general-purpose model from avoiding shortcomings of the electronic-structure reference.
  • Organic-crystal extension: Quantum sampling causes a large downward shift and broadening of 1H chemical-shielding distributions in α-succinic acid, especially for hydrogen-bonded protons.The three potentials show near-perfect agreement for these distributions under both classical and quantum sampling.
  • Functional properties: For barium titanate, PET-MAD and bespoke models agree on phase-transition temperatures within 30 K at the largest discrepancy, despite all transitions being underestimated relative to experiment.All models also identify the distinct phases and reproduce the qualitative dielectric-response increase approaching the ferroelectric transition.

III. DISCUSSION

PET-MAD combines a diverse, internally consistent dataset with a lightweight transformer-based architecture to provide broadly applicable atomistic simulations. Its compact training set and architecture support accuracy, speed, uncertainty estimation, and targeted fine-tuning.

  • Discussion: PET-MAD uses a dataset designed around internal DFT consistency and broad structural and chemical diversity.These principles support a coherent structure–energy mapping across the modeled systems.
  • Discussion: LLPR provides low-cost uncertainty estimates and supports propagating prediction uncertainty through difficult workflows such as molecular dynamics.The method uses last-layer feature covariance and can support thermodynamic uncertainty estimates.
  • Discussion: The dataset construction achieved >95% convergence for most subsets, with MC3D-random as the exception at about 55%.The lower convergence of MC3D-random is attributed to its extremely non-equilibrium structures.
  • Discussion: The model employs the Point Edge Transformer, a graph neural network whose message-passing layers use arbitrarily deep transformers.Geometric information and chemical species are incorporated into bond-level representations that are aggregated into target properties.
  • Discussion: PET imposes no explicit rotational symmetry constraint but learns approximate invariance through data augmentation while retaining high expressivity and computational speed.Its single-layer formulation can represent high body order and angular resolution.

D. Training of PET-MAD

PET-MAD is selected through accuracy–speed optimization and can be adapted efficiently to specialized systems. LoRA fine-tuning improves low-data accuracy while retaining useful behavior on generic structures.

  • Training of PET-MAD: LoRA fine-tuning adds trainable low-rank matrices to attention blocks to improve specialized-task performance while mitigating catastrophic forgetting.The standard configuration uses rank 8 and scaling parameter 0.5.
  • Training of PET-MAD: Fine-tuning is always beneficial in the low-data regime compared with training a specialized model from scratch.On larger datasets, specialized models can sometimes exceed fine-tuned models, but not for several investigated systems.
  • Training of PET-MAD: LoRA-finetuned models retain varying accuracy on generic MAD structures while producing practically equivalent observables to fully specialized models.The authors therefore recommend LoRA for adapting PET-MAD to a specific application.
  • Training of PET-MAD: The architecture search varied cutoff radius, message-passing layers, transformer layers, hidden dimensionality, and attention heads.Each hyperparameter combination was trained and evaluated for validation accuracy and inference time.
  • Training of PET-MAD: The Pareto-selected PET-MAD architecture uses a 4.5 Å cutoff, 2 message-passing layers, and 2 transformer layers per message-passing layer.The selection balances validation accuracy with inference time on a single NVIDIA GH200 GPU.

S2. DETAILS OF BENCHMARKING SUBSETS SELECTION

Benchmark interpretation depends strongly on matching DFT settings between training and evaluation data. PET-MAD’s apparent error on Matbench Discovery largely reflects inconsistent reference energetics rather than model accuracy alone.

  • Benchmark subset selection: The benchmark suite compares PET-MAD with four universal MLIPs across MPtrj, Matbench Discovery, Alexandria, SPICE, MD22, and OC2020 S2EF.Subsets were recalculated with DFT settings compatible with each model’s training data.
  • Benchmark subset selection: DFT inconsistency can introduce systematic prediction errors of up to 30 meV/atom.The authors emphasize consistency both during training and when assessing benchmark accuracy.
  • Benchmark subset selection: 138 meV/atom falls to 41 eV/atom for PET-MAD when WBM energies are compared with a consistent MAD DFT reference rather than the default WBM values.The reported 138 meV/atom comparison is dominated by a 120 meV/atom discrepancy in the underlying DFT reference.
  • Benchmark subset selection: PET-MAD was excluded from the Matbench Discovery model list because inconsistent baseline DFT data prevent representative accuracy comparisons.The phonon-properties portion was also omitted from this discussion because it was studied separately.
  • Benchmark subset selection: The MAD dataset uses uniform high plane-wave and charge-density cutoffs to maintain internally consistent energetics across systems.For some elements, isolated-atom energy differences between MAD and recommended SSSP settings reach 10–100 meV/atom.
  • Benchmark subset selection: MAD includes short-distance configurations that produce diatomic curves with quantitatively accurate repulsive behavior and avoid artificial dimers.The dimer evaluation spans distances from 0.9 to 5.0 times the corresponding covalent radius.

S6. GEOMETRY OPTIMIZATION WITH PET-MAD

The supplementary analyses examine PET-MAD’s optimization behavior, uncertainty estimates, phonon predictions, and retained accuracy after fine-tuning. They connect model uncertainty with deviations in derived observables and compare fine-tuned models with specialized baselines.

  • Geometry optimization: Geometry optimization statistics were collected from 1000 LBFGS runs for wurtzite BeO, zincblende BeTe, and rocksalt LiBr.The workflow used ASE calculators, random structural displacements, and a maximum force threshold of 0.05 eV/Å.
  • Uncertainty quantification: LLPR estimates prediction uncertainty from last-layer feature covariance at nearly no additional cost compared with raw predictions.Its sampled last-layer ensembles can propagate uncertainty through molecular dynamics and thermodynamic averages.
  • Geometry optimization: PET-MAD’s timing distributions include incomplete optimization runs as a separate count from the Gaussian-approximated timing band.The timing analysis covers 1000 geometry optimization runs.
  • Phonon calculations: For BeO, optical phonon bands are slightly softer than the reference while remaining consistent with MAD-settings calculations at an RMSD of 3 cm^-1.The increased ensemble variance for optical modes reflects this deviation in the uncertainty quantification.
  • Fine-tuning: Fine-tuned models are compared with bespoke models from scratch and retain some accuracy on the base MAD dataset while matching specialized observables.The comparison reports energy and force errors on the MAD test set for each fine-tuned model.

A. PET-MAD

PET-MAD maintains useful accuracy across diverse datasets and can be adapted through fine-tuning. Its lightweight model is competitive in several settings, although learning curves indicate further training data could improve performance.

  • Learning curves: 134.7 meV/Å force MAE is achieved on the MAD test set using only 20% of the training data.Increasing the training fraction gradually lowers test errors for energies and forces.
  • Model capacity: The lightweight PET-MAD model shows no saturation in the test-set learning curve, indicating that additional training data could still improve accuracy.Its compact design is described as suited to avoiding overfitting, while greater expressiveness could be obtained by adding GNN and transformer layers.
  • Fine-tuning: Fine-tuned PET-MAD models show dataset-dependent advantages over training from scratch for Li3PS4, HEA25S, water, and BTO.The supplied comparisons include LoRA and full fine-tuning, with outcomes varying by dataset and training regime.
  • Fine-tuning: For Li3PS4, complete fine-tuning outperforms the other scenarios by about 20%, while LoRA produces slightly higher force errors than the bespoke PET model.The small force-error difference does not translate into a large difference in ionic conductivity values.
  • Fine-tuning: HEA25S fine-tuned models outperform the bespoke model on the test set, likely because the training subset contains only 1975 structures versus about 30,000 originally.Fully fine-tuned models improve more as the training-data amount increases, while both fine-tuned models behave similarly in the low-data regime.
  • Fine-tuning: Water fine-tuning is advantageous through approximately 20% of the 1228-structure training set, after which higher LoRA rank partly remedies reduced accuracy.The comparison includes training from scratch and LoRA ranks 8 and 32.

F. Quantum nuclear effects in NMR crystallography

The section combines bespoke auxiliary modeling with simulation protocols for evaluating NMR-related properties and functional responses. The supplied passages emphasize model construction, fine-tuning comparisons, and transport calculations based on charge fluxes.

  • NMR crystallography: Chemical shieldings in succinic acid crystals are predicted with one SOAP-based linear model per central species and evaluated in a parity plot.The auxiliary model uses SOAP descriptors computed with the featomic library.
  • Model adaptation: Fine-tuned models outperform training from scratch for the BTO learning curves, while LoRA rank 8 gives slightly smaller force errors and updates fewer parameters.The supplied comparison also states that LoRA is faster than full fine-tuning.
  • Transport methodology: Ionic conductivities are computed with Green–Kubo linear-response theory, using charge-flux correlations along molecular-dynamics trajectories.The charge flux depends on atomic velocities and charges, while the time evolution and thermal averaging enter the transport calculation.
  • Water comparison: The water learning curves compare training from scratch with LoRA fine-tuning at ranks 8 and 32 against base PET-MAD prediction error.The supplied figure caption identifies the compared training scenarios but does not report numerical outcomes.

B. Melting point of GaAs

The GaAs melting-point workflow uses interface pinning to balance liquid and solid phases and identifies the transition where the average pinning force vanishes. PET-MAD uncertainty is propagated into the melting-point estimate through LLPR-based reweighting.

  • Interface pinning: GaAs melting points are computed with interface-pinning simulations of liquid–solid slabs, where equal phase chemical potentials define coexistence.The setup constrains coexisting phases separated by a planar interface.
  • Supplementary comparisons: The supplied supplementary figures include learning curves for succinic acid and BTO, but the passages provide no numerical melting-point comparison for GaAs.The BTO figure compares training from scratch, LoRA rank 8, and full fine-tuning.
  • Transition determination: The melting point is located by root finding the temperature at which the average pinning force is zero.Simulations use PLUMED for the collective variable and LAMMPS for molecular dynamics.
  • Simulation protocol: The interface simulations use 1152-atom GaAs supercells and production trajectories of 1 ns with a 4 fs timestep over 950–1200 K.Temperatures are sampled in 50 K intervals, with pressure controlled near 1 bar at the interface.
  • Uncertainty quantification: PET-MAD energy uncertainties are propagated to melting-point uncertainties through thermodynamic reweighting of interface-pinning observables.LLPR supplies a committee of 128 potential-energy predictions at each simulation timestep.

C. Surface segregation in high-entropy alloys

The surface-segregation analysis uses replica-exchange molecular dynamics with Monte Carlo swaps to relax a CoCrFeMnNi (111) slab. Segregation is quantified with Gibbs surface excess, while related simulation protocols address quantum and dielectric observables.

  • Surface segregation: CoCrFeMnNi surface relaxation uses a 539-atom (111) slab, 16 replicas, Monte Carlo atom swaps, and 200 ps NPT trajectories from 500 to 1200 K.Both pre-trained and fine-tuned PET-MAD potentials are used in identical REMD/MC runs.
  • Surface segregation: Gibbs surface excess Γa measures each element’s surface segregation relative to the bulk, with Γa > 0 indicating enrichment and Γa < 0 indicating depletion.The bulk region is defined as a 10 Å-thick region around the slab center.
  • Quantum nuclear effects: PIMD represents nuclear quantum effects with P equivalent replicas at temperature PT coupled by harmonic springs between corresponding atoms.The treatment becomes exact as P → ∞ and reduces to classical molecular dynamics when P = 1.
  • Phase transitions: BTO phase transitions are identified by clustering sampled structures and locating the temperature where the relative chemical potential between phases is zero.The transition temperature is obtained by linearly fitting the chemical-potential difference as a function of temperature.
  • Dielectric response: The dielectric tensor is computed from cell-dipole covariance, becoming scalar in the centrosymmetric cubic phase and anisotropic after transition to the tetragonal phase.Parallel and perpendicular components are obtained by projecting onto the dipole direction and its orthogonal complement.

S12. NON-CONSERVATIVE MD

PET-MAD simulations compare conservative forces, faster direct-force predictions, and multiple time stepping in molten BMIM-Cl. Direct forces produce unphysical behavior, while multiple time stepping retains stable trajectories and most of the computational advantage.

  • Methods: Direct-force predictions are approximately 2 times faster than conservative forces, while multiple time stepping is approximately 1.8 times faster than conservative MD.Multiple time stepping evaluates conservative forces every 8 steps to correct the non-conservative trajectory.
  • Constant-energy simulations: In constant-energy simulations, direct forces cause rapid kinetic-energy drift, molecular dissociation, and unphysical outcomes.The simulations use a 0.5 fs time step and a velocity Verlet integrator.
  • Constant-energy simulations: Multiple time stepping yields a stable trajectory entirely consistent with conservative molecular dynamics.It avoids the catastrophic drift observed with non-conservative forces.
  • Structural and dynamic properties: Multiple time stepping reproduces conservative calculations for both the Cl-Cl pair correlation function and Cl diffusion within statistical uncertainty.Direct-force trajectories instead produce extremely fast Cl diffusion, consistent with their much higher steady-state temperature.
  • Thermostatted simulations: An aggressive global thermostat can avoid overall temperature drift but breaks energy conservation and produces different steady-state kinetic energies across degrees of freedom.Multiple time stepping avoids these thermostat-related artifacts while retaining most of the computational advantage.
Loading 2503.14118v2…