Source-linked AI summary
Machine learning force fields and coarse-grained variables in molecular dynamics: application to materials and biological systems
Paraskevi Gkeka, Gabriel Stoltz, Amir Barati Farimani, Zineb Belkacemi, Michele Ceriotti, John Chodera, Aaron R. Dinner, Andrew Ferguson, Jean-Bernard Maillet, Hervé Minoux, Christine Peter, Fabio Pietrucci, Ana Silveira, Alexandre Tkatchenko, Zofia Trstanova, Rafal Wiewiora, Tony Leliévre
TL;DR
The review addresses how machine learning can extract useful representations and models from molecular-dynamics and ab-initio data for atomistic systems. It surveys force-field construction, reaction-coordinate and collective-variable identification, enhanced sampling, and biological applications, concluding that these methods support accurate and versatile simulations while their performance depends on data coverage, representation, regression, and physical knowledge.
Problem
Atomistic simulations generate extensive data, while empirical force fields and reaction coordinates require accurate, transferable representations for free-energy calculations and enhanced sampling.
Method
The review synthesizes machine-learning approaches for fitting force fields and quantum-chemical observables, identifying collective variables, constructing effective models, and analyzing biological systems.
Results
Machine-learning force fields and collective-variable methods have supported accurate simulations of molecules and materials and applications to biological complexity and drug discovery.
Takeaways & Limitations
The paper highlights machine learning as a versatile framework for combining atomistic data, physical knowledge, and reduced representations in molecular-dynamics studies.
Abstract
from arXiv · showhide
Machine learning encompasses a set of tools and algorithms which are now becoming popular in almost all scientific and technological fields. This is true for molecular dynamics as well, where machine learning offers promises of extracting valuable information from the enormous amounts of data generated by simulation of complex systems. We provide here a review of our current understanding of goals, benefits, and limitations of machine learning techniques for computational studies on atomistic systems, focusing on the construction of empirical force fields from ab-initio databases and the determination of reaction coordinates for free energy computation and enhanced sampling.
1. Introduction
The review examines machine learning for constructing force fields, defining reaction coordinates, and analyzing biological systems and drug discovery. It emphasizes systematic, data-driven representations for molecular dynamics and enhanced sampling.
- Motivation: Coarse-grained representations simplify atomistic systems and can serve as surrogate models for enhanced sampling.They may describe molecular conformational changes or material phase transitions using selected structural variables.
- Motivation: Machine learning is increasingly used to automate reaction-coordinate and atomic-descriptor construction beyond empirical methods and chemical intuition.The review considers both supervised and unsupervised approaches.
- Review scope: For force fields, machine learning fits potentials or quantum-chemical observables from databases, with accuracy, transferability, database design, regression choice, and physical knowledge treated as key factors.The review also discusses surrogate models for quantities such as NMR chemical shieldings and IR dipole moments.
- Review scope: For biological systems, the review covers clustering, Markov state models, efficient collective variables, enhanced sampling, and drug-discovery applications.These examples are presented as ways to study thermodynamic and mechanistic behavior at molecular resolution.
2. Machine learning force fields and Potential of Mean Force
This section reviews machine-learning force fields, potential-of-mean-force models, collective-variable identification, and biological applications. It emphasizes database coverage and data-driven effective models for atomistic and coarse-grained simulations.
- Force fields: Machine learning can approximate quantum-mechanical forces, energies, and other observables while reducing the computational cost of ab-initio calculations.The review frames this as a route to force fields with near ab-initio accuracy and atom-number-based scaling.
- Collective variables: The review addresses collective-variable identification from databases covering either the full configuration space or a restricted metastable state.It also considers learning effective free-energy and dynamical models along the selected coordinate.
- Collective variables: Effective models along a reaction coordinate may approximate free energies or drift, diffusion, metric, and memory terms.The latter can be constructed through projections à la Mori-Zwanzig.
- Biological applications: Machine-learning applications in biology and drug discovery use molecular-dynamics data to study biological complexity at molecular resolution.The review presents these applications as part of its discussion of real-world uses.
- Database design: Force-field accuracy and transferability depend strongly on database coverage of the thermodynamic conditions where the potential will be used.Active learning can add configurations when two neural networks disagree substantially on a new configuration.
A. Setting up a database.
Setting up a machine-learning force-field database requires representations that respect physical symmetries and a coordinated choice of descriptors and regression models. The database and representation determine how configurations are presented for fitting.
- Representations: Cartesian atomic coordinates are generally unsuitable as machine-learning inputs because target properties are invariant to atom permutations, translations, rotations, and reflections.Feature mappings are therefore designed to satisfy these symmetry requirements.
- Fitting pipeline: The fitting pipeline separates atomic-configuration descriptors from the subsequent regression used to determine model parameters.Descriptor complexity and regression complexity can be traded against one another.
- Fitting pipeline: Simple physically motivated descriptors may be paired with complex regressors, whereas more informative descriptors can support simpler regression models.This design choice connects representation quality with model complexity.
B. Descriptors and regression methods.
After choosing an atomic descriptor, the regression method must be selected according to the system and the desired optimization and uncertainty properties. Neural-network, kernel, and bilinear approaches offer different computational characteristics.
- Regression choice: Regression-method choice is crucial after descriptor selection and depends strongly on the system under study.The review distinguishes neural networks from kernel-based and bilinear methods.
- Optimization: Neural-network training involves high-dimensional non-convex optimization, while kernel and bilinear methods can yield better-behaved optimization problems.Kernel and bilinear objectives may be solved analytically through matrix inversion.
- Uncertainty: Kernel methods can provide variance-based prediction uncertainties, whereas error quantification is harder with neural networks.Thus, regression choice affects both fitting behavior and available error estimators.
B.2. Choosing the regression method.
Regression methods for force-field learning span neural networks, kernel approaches, and deep-network models, with choices shaped by the system and available reference data.
- B.2. Choosing the regression method.: Deep-network potential-energy-surface methods can learn similarity measures directly from training data without an a priori similarity definition.
- B.2. Choosing the regression method.: Organic-molecule force fields face stricter reference-data constraints because coupled-cluster CCSD(T) calculations are expensive and feasible for only hundreds of calculations even for simple molecules.DFT is often considered sufficiently accurate for solids but not for organic molecules.
- B.2. Choosing the regression method.: Behler–Parrinello neural networks and kernel-based GAP models achieve 1-2 meV/atom accuracy for several solids.Reported examples include C, Si, Cu, and TiO2.
C. Synergy between physics, chemistry, mathematics and ML approaches.
The review emphasizes combining machine learning with physical structure and mathematical constraints to improve force-field simulations while exposing limitations from long-range interactions and model choices.
- C. Synergy between physics, chemistry, mathematics and ML approaches.: ML force fields incorporate translational, rotational, and permutational symmetries, and energy–force learning can enforce exact energy conservation.The force is learned as the negative gradient of the energy.
- C. Synergy between physics, chemistry, mathematics and ML approaches.: Table 1 summarizes key machine-learning methods developed for force-field development.
- C. Synergy between physics, chemistry, mathematics and ML approaches.: Long-range electrostatic and plasmon-like interactions can extend 20-30 nanometers, whereas many ML models cut interactions off at 5-6 Å.The review therefore identifies coupling short-range ML models with explicit non-covalent-interaction physics as necessary.
- C. Synergy between physics, chemistry, mathematics and ML approaches.: A simpler Ziegler–Biersack–Littmark starting potential yields better results than fitting the small, noisy, rugged difference from a good classical potential to an ab-initio potential.
D. Perspectives for ML approaches to the determination of force fields.
The review identifies preconditioning, multi-objective optimization, data sensitivity, numerical stability, and intensive-property prediction as open directions for machine-learning force fields.
- D. Perspectives for ML approaches to the determination of force fields.: ML can learn the difference between an acceptable empirical force field and a DFT model as a form of preconditioning.Kernel methods have been used to build potentials on top of pre-existing two-body and three-body classical potentials.
- D. Perspectives for ML approaches to the determination of force fields.: Potential optimization depends on arbitrary weighting of energy, force, stress, and sometimes bond-distance terms, motivating a unified cost-function definition.Different weights can produce infinitely many optimal potentials.
- D. Perspectives for ML approaches to the determination of force fields.: The learned parameters can be sensitive to the data, including the fraction of elements assigned to training versus testing.
- D. Perspectives for ML approaches to the determination of force fields.: The numerical stability of machine-learning potentials for time integration remains an open theoretical question, although preliminary results suggest they may be smoother than empirical potentials.
- D. Perspectives for ML approaches to the determination of force fields.: Predicting intensive rather than extensive properties remains very challenging for reasons that are not yet understood.
E. Bottom-up coarse-graining force fields: From PES to FES.
Bottom-up coarse-graining reduces dimensionality by grouping atoms and seeks effective interactions that reproduce atomistic configurational sampling. Machine learning can represent the resulting many-body potential of mean force using local descriptors and regression.
- Coarse-grained models group several atoms to reduce the dimensionality of classical phase space.
- Bottom-up strategies determine effective Hamiltonians that reproduce mapped atomistic configurational sampling.
- Gaussian process regression can describe many-body coarse-grained potentials of mean force through local multibody terms and descriptors.
- Free energies and potentials of mean force must be inferred from sampled probability densities or mean forces rather than obtained directly from molecular dynamics.
- Collective-variable identification is treated alongside coarse-grained modeling as a route to reduced representations for enhanced sampling.
B. Data-driven discovery of high-variance and slow collective variables.
Data-driven collective-variable discovery targets high-variance or slowly evolving molecular degrees of freedom that are difficult to identify by intuition. Linear and nonlinear dimensionality-reduction methods can then support enhanced sampling and dynamical analysis.
- Emergent collective variables are difficult to intuit, motivating systematic estimation from molecular simulation data.
- High-variance methods include PCA and nonlinear manifold-learning techniques that project molecular data into low-dimensional collective-variable spaces.
- Enhanced-sampling protocols can use learned collective variables from partial sampling to leave metastable states and explore new configurations.
- Diffusion maps identify slowly evolving principal modes by approximating a Fokker–Planck operator from trajectory point clouds.
- Diffusion maps can support committor computation in high dimensions, while their low computational complexity aids molecular-trajectory analysis.
- Projected low-dimensional dynamics are generally non-Markovian unless a time-scale separation is present.
D. Extracting dynamical information from trajectory data.
Trajectory-based methods use identified collective variables or metastable states to extract dynamical information from short molecular-dynamics simulations. In a dipeptide example, approximately 100 short trajectories encoded profiles needed for thermodynamic and kinetic modeling.
- Galerkin projection estimates dynamical statistics by expanding them in basis functions and fitting generator matrix elements from short molecular-dynamics trajectories.
- ∼100 short trajectories of a few picoseconds encoded information needed to reconstruct free-energy, friction, and mass profiles in a dipeptide test case.
- The approach was presented as suitable for high barriers and non-Markovian dynamics while providing thermodynamics and kinetics using unbiased molecular dynamics.
- Machine-learning methods are discussed for protein conformational analysis and drug-discovery applications involving metastable states and binding mechanisms.
- Kinetically motivated dimensionality reduction and cross-validation helped quantify experimentally validated ensembles of kinetically relevant protein macrostates.
B.1. Conformational-specific targeting of proteins using cryptic binding sites.
Machine-learning methods can expose protein conformations and binding sites that are not apparent from known crystallographic structures. Applied to trajectories and conformational landscapes, these methods support cryptic and allosteric drug-discovery strategies.
- Cryptic and allosteric sites: Non-orthosteric binding sites can reveal alternative protein conformations and may improve inhibitor selectivity relative to orthosteric targeting.The review cites allosteric inhibitors and transient sites that are absent from known crystallographic conformations.
- Conformational landscapes: The SETD8 workflow combines crystallographically derived structural chimeras, parallel molecular dynamics, and Markov state models to construct apo- and SAM-bound conformational landscapes.The figure presents structural-chimera construction and a dynamic conformational-landscape workflow based on MSM.
- Functional consequences: Subtle protein-conformation changes can distort TNFα trimer assembly and affect downstream TNFR1 signaling, motivating compounds that stabilize the altered trimer.This example connects conformational analysis with the design of compounds for TNFα-related diseases.
- Learning binding-site coordinates: ML algorithms such as k-means and Markov models reduce the dimensionality of drug-binding events to help identify allosteric binding sites.tICA can learn reaction coordinates from drug-binding and unbinding trajectories generated with different initial seeds.
- Learning conformational changes: Variational autoencoders and tICA can identify subtle, sequential protein motions involved in activation pathways such as those of GPCRs.These methods address the difficulty of finding relevant motions in a high-dimensional conformational space.
5. Concluding remarks and perspective
The review organizes machine learning applications in molecular simulation around modeling, numerical efficiency, and data analysis. It also emphasizes shared databases, benchmarks, maintained software, and explicit testing of transferability as priorities for the field.
- Global objectives: The review identifies three objectives for coarse-graining: improving models, accelerating numerical methods, and analyzing simulation data.Examples include better force fields and energy surfaces, collective variables for enhanced sampling, and Markov state models for identifying states.
- Global objectives: Machine learning can support modeling through improved force fields and potential energy surfaces, numerical efficiency through collective variables, and data analysis through Markov state models.These applications span ab-initio-based potentials, enhanced sampling, and interpretation of molecular-dynamics trajectories.
- Community infrastructure: The review recommends established or shared databases, standard maintained software packages, common benchmarks, and prediction contests with broadly accessible participation.These practices are presented as ways to coordinate goals and priorities across approaches.
- Transferability: Transferability should be emphasized by training with databases and testing on different databases rather than evaluating only within the original data domain.The proposed direction focuses evaluation on performance beyond the database used for model development.