Source-linked AI summary

Forces are not Enough: Benchmark and Critical Evaluation for Machine Learning Force Fields with Molecular Simulations

Xiang Fu, Zhenghao Wu, Wujie Wang, Tian Xie, Sinan Keten, Rafael Gomez-Bombarelli, Tommi Jaakkola

arXiv:2210.07237v2physics.comp-phcs.LGphysics.chem-ph

TL;DR

ML force fields are commonly judged by force and energy errors, although their practical purpose is to generate realistic MD trajectories. This paper introduces a diverse simulation benchmark with task-specific metrics, showing that force accuracy can diverge from simulation quality and that stability remains a key challenge.

  • Problem

    Existing ML force-field benchmarks primarily measure force and energy prediction rather than realistic MD trajectories and system observables.

  • Method

    The paper builds a benchmark suite spanning diverse MD systems, simulation protocols, and quantitative metrics, and evaluates SOTA ML force-field models.

  • Results

    Many existing models perform inadequately on simulation-based benchmarks despite accurate force prediction, with failures arising in stability, sampling, and observable recovery.

  • Takeaways & Limitations

    Simulation-based metrics, especially stability, should complement force accuracy when assessing the practical utility of ML force fields.

  • Takeaways & Limitations

    Performance is highly case-dependent, and challenging systems may require more expressive atomistic representations or broader training coverage.

Abstract

from arXiv · show

Molecular dynamics (MD) simulation techniques are widely used for various natural science applications. Increasingly, machine learning (ML) force field (FF) models begin to replace ab-initio simulations by predicting forces directly from atomic structures. Despite significant progress in this area, such techniques are primarily benchmarked by their force/energy prediction errors, even though the practical use case would be to produce realistic MD trajectories. We aim to fill this gap by introducing a novel benchmark suite for learned MD simulation. We curate representative MD systems, including water, organic molecules, a peptide, and materials, and design evaluation metrics corresponding to the scientific objectives of respective systems. We benchmark a collection of state-of-the-art (SOTA) ML FF models and illustrate, in particular, how the commonly benchmarked force accuracy is not well aligned with relevant simulation metrics. We demonstrate when and how selected SOTA methods fail, along with offering directions for further improvement. Specifically, we identify stability as a key metric for ML models to improve. Our benchmark suite comes with a comprehensive open-source codebase for training and simulation with ML FFs to facilitate future work.

1 Introduction

The paper argues that ML force fields should be evaluated through realistic MD simulations rather than force accuracy alone. It introduces a diverse benchmark with simulation protocols and metrics, and examines model failures and stability.

  • Motivation: MD simulations provide atomistic insights, but quantum-mechanical potential-energy calculations are computationally expensive.ML force fields aim to accelerate simulations while retaining quantum-chemical accuracy.
  • Motivation: Existing ML force-field evaluations often use force and energy prediction accuracy without testing whether models produce useful simulations.The paper identifies this mismatch as a central gap in current benchmarking.
  • Contributions: The benchmark evaluates diverse MD systems with simulation protocols and quantitative metrics tailored to their scientific objectives.The design considers system diversity, computational cost, and established simulation practice.
  • Findings: Many existing models are inadequate on simulation-based benchmarks even when their force predictions are accurate.Figure 1 illustrates that force error does not necessarily align with stability or structural-distribution metrics.
  • Contributions: The authors summarize common simulation failure modes and discuss their causes and potential solutions.An open-source codebase for training and simulation is provided to facilitate future work.

2 Preliminaries

The preliminaries describe ML force-field learning, Newtonian MD integration, and observable-based evaluation. Simulation observables connect trajectories to equilibrium distributions and macroscopic material properties.

  • Training: An ML force field learns a potential energy surface from atomic coordinates by fitting forces and energies in a training dataset.The learned surface maps x ∈ R^(N×3) to an energy and supports force evaluation.
  • MD simulation: MD simulations integrate Newtonian equations of motion using forces obtained by differentiating the learned potential energy surface.Thermostats and barostats augment the dynamics to impose system- and task-dependent thermodynamic conditions.
  • Observables: Simulation trajectories produce observables that characterize system states at different granularities.Examples include RDFs, virial stress tensors, mean-squared displacement, and dihedral angles.
  • Observables: Under the ergodic hypothesis, time averages of observables converge to distributional averages under the Gibbs measure.The observables require the system to reach equilibrium before they can be calculated.
  • Evaluation: The benchmark proposes metrics based on established observables for the respective system types.These metrics are intended to connect ML-driven simulations with experimental measurements and macroscopic properties.

3 Related Work

Prior ML force-field benchmarks mainly assess force or energy prediction, often on small molecules, while newer studies address relaxation or individual simulation systems. The paper targets broader, quantitative comparison across diverse MD tasks.

  • ML force fields: ML force fields use expressive regressors, including kernel methods and neural networks on symmetry-preserving atomic representations.Graph neural network architectures have recently become increasingly popular.
  • Existing benchmarks: Existing benchmarks mostly focus on force and energy prediction, with small molecules as typical systems.OC20 and OC22 instead emphasize structural relaxation, whose goal is a final relaxed structure or energy.
  • Existing benchmarks: Structural relaxation and force-prediction tasks do not characterize system properties under a structural ensemble.This limits their coverage of simulation observables relevant to MD behavior.
  • Benchmark scope: The benchmark visualizes organic molecules, water systems, alanine dipeptide, and LiPS as representative simulation systems.These systems span molecular, liquid, peptide, and crystalline-material settings.
  • Research gap: Previous simulation studies often focus on a single system and model, while protocols and quantitative comparison metrics remain debated.Results are frequently presented qualitatively or as figures, making scalar performance comparison difficult.

4 Datasets

The benchmark covers gas-phase organic molecules, water, alanine dipeptide, and the crystalline conductor LiPS. Each system is paired with simulation metrics that probe stability, structural distributions, dynamics, or conformational sampling.

  • Dataset scope: The selected systems target complex interatomic interactions and simulation observables beyond gas-phase small-molecule force prediction.The benchmark focuses on atomic-level MD with intermolecular interactions across multiple scales.
  • MD17 molecules: Four MD17 molecules are evaluated using force error, stability, and interatomic-distance distributions.Aspirin, ethanol, naphthalene, and salicylic acid are each simulated for five 300-ps trajectories at 500 K.
  • Water: Water is evaluated through force error, stability, element-conditioned RDFs, and liquid diffusion coefficients.The dataset contains 100,000 equilibrium structures sampled from a 1-ns trajectory at 300 K, with multiple training-set sizes.
  • Alanine dipeptide: Alanine dipeptide is assessed through conformational sampling of its metastable states using the central ϕ and ψ dihedral angles.The reference data use explicit-water simulations, while the learned force field uses implicit solvation.
  • LiPS: LiPS is evaluated on force error, stability, RDF recovery, and Li-ion diffusivity as a representative crystalline superionic conductor.The dataset contains 25,000 structures, with 19,000 used for training and 1,000 for validation.

5 Evaluation Metrics

The benchmark evaluates learned MD simulations through physical observables and stability, rather than relying only on force accuracy. Metrics cover structural, dynamical, thermodynamic, and stability behavior across relevant molecular systems.

  • Observable-based evaluation: Simulation observables connect learned MD trajectories to experimentally measurable structure and macroscopic properties.The benchmark emphasizes observables such as RDFs, diffusivity, and free-energy surfaces because accurate recovery supports more sophisticated analyses.
  • Structural metrics: RDF error compares reference and model-predicted equilibrium radial distributions by integrating their absolute difference over distance.RDFs describe how density varies with distance from a particle and characterize structural and thermodynamic properties.
  • Dynamical metrics: Diffusivity is estimated from mean-squared displacement and averaged across valid trajectories, requiring stable runs of at least 100 ps for water and 40 ps for LiPS.The tracked particles are all 64 oxygen atoms for water and all 27 lithium ions for LiPS.
  • Thermodynamic metrics: Alanine dipeptide free-energy accuracy is measured by integrating absolute reconstruction error along the physically informative dihedral coordinates ϕ and ψ.The free-energy surface is computed from configuration probabilities using F(ξ) = −kBT ln p(ξ).
  • Stability metrics: Stability is monitored through deviations in equilibrium statistics, with RDF-based criteria for periodic systems and bond-length criteria for flexible molecules.For water, instability occurs when any element-conditioned RDF exceeds the threshold; thresholds are τ = 1 ps and ∆ = 3.0 for water, and τ = 1 ps and ∆ = 1.0 for LiPS.
  • Stability metrics: Observable metrics are computed only over stable trajectory segments to separate prediction accuracy from catastrophic simulation failure.The thresholds are intentionally relaxed to detect failures that the model cannot recover from.

6 Experiments

Experiments across organic molecules, water, alanine dipeptide, and LiPS show that force accuracy does not reliably predict simulation quality. Stability is a key practical bottleneck, while model architecture and dataset size can change the trade-off between accuracy, stability, observables, and speed.

  • Force prediction is insufficient for evaluating ML force fields because it generally does not align with simulation stability or ensemble-property estimation.
  • MD17: On MD17, SphereNet and GemNet-T/dT often achieve low force error but collapse before completion, whereas DeepPot-SE performs well on simulation metrics despite relatively high force error.
  • MD17 and Water: NequIP performs best overall on MD17 and water, while DeepPot-SE matches NequIP on water-90k and is more than 20 times faster.
  • Water: On water, GemNet-T/dT and DimeNet rank among the best in force prediction but lack stability, while DeepPot-SE shows decent stability and accurate simulation statistics.
  • Water: More training data usually reduces force error, but does not necessarily improve stability or ensemble-statistics estimation.
  • Alanine dipeptide: Alanine dipeptide exposes severe stability challenges, with only GemNet-T and NequIP completing at least one 5 ns simulation and both producing inaccurate statistics.
  • LiPS: For LiPS, GemNet-T and GemNet-dT combine excellent force prediction, stability, and observable recovery, with GemNet-dT running 2.6 times faster.

7 Failure Modes: Causes and Future Directions

The paper traces simulation failures to poor exploration of configuration space, non-conservative forces, and instability that persists despite accurate force predictions or more data. It proposes broader training coverage and uncertainty-driven data acquisition as possible remedies.

  • Failure causes: NequIP and GemNet-T can fail to reconstruct alanine-dipeptide free-energy surfaces because they undersample transition regions and parts of configuration space.NequIP can also become unstable from a low-density metastable state, producing a free-energy surface that deviates from the reference.
  • Failure causes: Simulation collapse usually follows short-lived local errors, after which trajectories enter nonphysical regions and forces become increasingly pathological.The collapse examples involve NequIP on alanine dipeptide and GemNet-T on water; the figure tracks maximum force over time.
  • Failure causes: Non-conservative forces can break time-reversal symmetry and prevent equilibration, while energy conservation coincides with different stability outcomes across GemNet variants and tasks.GemNet-T conserves energy and is stable for alanine dipeptide, whereas GemNet-dT fails energy conservation and is unstable there but performs well on LiPS with a thermostat.
  • Failure causes: More training data reduces force error but does not consistently improve stability, especially for GemNet-T/dT and DimeNet.Reducing the simulation timestep for GemNet-T on water also fails to improve stability.
  • Future directions: Training on distorted and off-equilibrium geometries or using active learning may improve robustness by preventing trajectories from entering nonphysical regions.Noise-based denoising, empirical prior energies, and post-prediction refinement are additional approaches discussed for instability.

8 Conclusion and Outlook

The paper concludes that ML force fields should be evaluated through realistic MD simulations rather than force error alone. Its benchmark exposes case-dependent failures and motivates more expressive, scalable representations and further benchmark development.

  • Conclusion: The authors provide a diverse suite of MD simulation tasks and a thorough comparison of state-of-the-art ML force fields.The benchmark is intended to reveal failure modes and insights for improving ML-based MD simulation.
  • Conclusion: The benchmark shows that force error alone is insufficient for assessing ML force fields’ practical utility in MD simulations.Simulation-based metrics are needed to reflect whether models produce useful trajectories and observables.
  • Outlook: Model performance is highly case-dependent, and challenging systems may require more expressive atomistic representations.The paper cites non-local descriptors for long-range interactions and local equivariant representations where scalability is critical.
  • Outlook: New datasets and benchmarks can support future work on ML for MD simulations, alongside approaches such as coarse-grained MD for larger length and time scales.The paper also points to enhanced sampling and differentiable simulations as related directions.

A Dataset details

The benchmark combines publicly available MD17 and LiPS data with water and alanine-dipeptide systems, including implicit-solvent and metadynamics settings. It also compares ML force fields with a classical OPLS baseline for molecular structure statistics.

  • Datasets: MD17 and LiPS are public datasets derived from path-integral and ab-initio molecular dynamics, respectively.MD17 incorporates quantum effects through Feynman path integrals, while LiPS uses ab-initio simulations with PBE and projector augmented wave pseudopotentials.
  • Classical baseline: OPLS simulations provide a classical-force-field comparison for four MD17 molecules, with resulting h(r) errors much higher than those of most ML force fields.The comparison uses vacuum, NVT simulations at 500 K with a 1 femtosecond timestep.
  • Water: The water dataset uses flexible SPC/E-fw at T = 300 K and P = 1 atm, with parameters fitted to experimental bulk properties.The cited properties include self-diffusion and dielectric constants.
  • Alanine dipeptide: Alanine dipeptide is simulated in explicit water with AMBER-03 and characterized by six free-energy local minima.Six simulations are initialized for each model from the six minima.
  • Implicit solvation: The implicit-solvent task incorporates the explicit solvent environment into a learned force field to reduce the number of simulated particles.The explicit system contains 1164 water molecules, which add substantial computational cost.
  • Metadynamics: Metadynamics accelerates rare-event sampling by depositing Gaussians along ϕ and ψ to estimate the alanine-dipeptide free-energy surface.The evaluation deposits Gaussians every 1 ps with height h = 1.2 and sigma σ = 0.35.

B Experimental details

Experiments evaluate learned trajectories with relaxed stability thresholds, symmetry-aware models, recorded observables, and dataset-specific metrics. Results show that force accuracy, stability, ensemble statistics, and scalability can diverge substantially across models and system sizes.

  • Evaluation: Stability is flagged only after trajectories enter highly unrealistic configurations, using thresholds based on reference RDF and bond-length fluctuations.Unstable simulations cannot recover from catastrophic failure, and examples include SchNet on Aspirin and GemNet-T on water.
  • Model design: ML force fields are expected to preserve permutational invariance and E(3) symmetry, with forces equivariant under translations, rotations, and reflections.Energy is invariant, while forces transform equivariantly under the symmetry operations.
  • Evaluation: Evaluation records atomic positions, temperatures, energies, and kinetics, then compares system-specific observables such as RDFs, diffusivity, MSD, and dihedral angles.The framework simulates trajectories with prescribed thermostats and lengths.
  • Results: In 512-molecule water, NequIP and SphereNet remain stable for 150 ps, but SphereNet fails to reproduce correct ensemble properties and most other models cannot support diffusivity computation.DeepPot-SE stability drops significantly, plausibly because limited message passing restricts long-range interaction modeling.
  • Results: NequIP’s 4-radius cutoff lowers force-prediction performance but preserves trajectory statistics while improving computational efficiency.All tested NequIP sizes and cutoffs are highly stable with similarly strong simulation-based metrics.
  • Results: More data nearly always lowers force error, but stability and ensemble-statistics estimates do not necessarily improve.NequIP remains stable across dataset sizes, while DeepPot-SE and SchNet show stability gains with more data.
  • Results: SchNet training on salicylic acid reduces force error and improves stability over epochs, while a structure-based split makes most models slightly worse without changing ranking trends.These findings separate training progression effects from split-dependent performance changes.
  • Results: ForceNet produces inaccurate h(r) curves for Aspirin, SchNet is less accurate on Aspirin and Ethanol, and LiPS RDFs are accurate for all models except unstable DeepPot-SE and inaccurate ForceNet.Water RDF failures can reflect inaccurate interaction modeling rather than instability alone.
Loading 2210.07237v2…