Source-linked AI summary

Machine Learning of coarse-grained Molecular Dynamics Force Fields

Jiang Wang, Simon Olsson, Christoph Wehmeyer, Adria Perez, Nicholas E. Charron, Gianni de Fabritiis, Frank Noe, Cecilia Clementi

arXiv:1812.01736v3physics.comp-phcs.LGstat.ML

TL;DR

Atomistic simulations cannot efficiently cover large macromolecular systems on biological timescales, while coarse-grained models must preserve selected properties despite dimensionality reduction and difficult many-body effects. The paper reformulates force matching as supervised learning, introduces invariant CGnets for learning coarse-grained free energies, and uses statistical learning theory and cross-validation to assess models. CGnets reproduce complex free-energy landscapes more effectively than classical few-body models, including Chignolin folding and unfolding with only Cα atoms and no solvent.

  • Problem

    Large macromolecular complexes remain inaccessible to atomistic simulation on biological timescales, while coarse-grained models must preserve selected properties and often require difficult-to-model many-body terms.

  • Method

    The paper formulates force matching as supervised learning, decomposes its error into Bias, Variance, and Noise, ranks models by cross-validation, and introduces invariant CGnets that learn free-energy functions from forces.

  • Results

    CGnets systematically outperform classical few-body coarse-graining and approximate Chignolin’s folding/unfolding landscape, capturing all free-energy minima with only protein Cα atoms and no solvent.

  • Takeaways & Limitations

    CGnets can encode multi-body effects and complex solvation free energies in compact coarse-grained models while preserving physically relevant invariances and constraints.

  • Takeaways & Limitations

    The demonstrated CG model is designed ad hoc for one molecule and is not transferable to different systems; larger, more complex molecules may require additional physical-constraint terms.

Abstract

from arXiv · show

Atomistic or ab-initio molecular dynamics simulations are widely used to predict thermodynamics and kinetics and relate them to molecular structure. A common approach to go beyond the time- and length-scales accessible with such computationally expensive simulations is the definition of coarse-grained molecular models. Existing coarse-graining approaches define an effective interaction potential to match defined properties of high-resolution models or experimental data. In this paper, we reformulate coarse-graining as a supervised machine learning problem. We use statistical learning theory to decompose the coarse-graining error and cross-validation to select and compare the performance of different models. We introduce CGnets, a deep learning approach, that learns coarse-grained free energy functions and can be trained by a force matching scheme. CGnets maintain all physically relevant invariances and allow one to incorporate prior physics knowledge to avoid sampling of unphysical structures. We show that CGnets can capture all-atom explicit-solvent free energy surfaces with models using only a few coarse-grained beads and no solvent, while classical coarse-graining methods fail to capture crucial features of the free energy surface. Thus, CGnets are able to capture multi-body terms that emerge from the dimensionality reduction.

Introduction

The paper reframes coarse-graining as supervised learning to build predictive models at reduced resolution, addressing limitations in sampling, energy matching, and conventional few-body potentials. It introduces CGnets, which learn free-energy functions with force matching, physical invariances, and regularization, and demonstrates improved representation of complex free-energy landscapes.

  • Motivation: Large macromolecular complexes remain inaccessible to atomistic simulation on biological timescales, motivating simplified predictive models.Such models can complement increasingly accessible but incomplete experimental data and help interpret molecular dynamics.
  • Problem: Coarse-graining cannot directly match atomistic potential energies after dimensionality reduction, so the properties to preserve must be defined explicitly.Optimal coarse-grained potentials often contain many-body terms that are difficult to represent with conventional energy functions.
  • Machine-learning formulation: The paper formulates force matching as supervised learning, decomposes error into Bias, Variance, and Noise, and uses cross-validation to rank models.This reframes coarse-graining from parameter fitting toward minimizing prediction error on unseen data.
  • CGnets: CGnets learn coarse-grained free energies whose coordinate gradients produce mean forces, while enforcing rotational and translational invariances.Because the coarse-grained free energy is initially unknown, training uses force information rather than directly fitted energies.
  • CGnets: Regularized CGnets add a prior energy that diverges for unphysical states, producing restoring forces without compromising accuracy within the training-data region.This addresses catastrophically wrong force predictions for configurations outside the sampled training region.
  • Results: CGnets outperform classical few-body coarse-graining on explicit-solvent systems, including Chignolin folding and unfolding represented using only protein Cα atoms and no solvent.For Chignolin, CGnets capture all free-energy minima, whereas the classical model cannot reproduce the folding/unfolding dynamics.

Theory and methods

The paper formulates coarse-graining as supervised learning: a lower-dimensional model is trained by force matching to reproduce mean forces while preserving thermodynamic consistency. CGnets learn conservative free energies, and regularization incorporates physical constraints to avoid unstable predictions outside the training data.

  • Coarse-grained representation: Coarse-graining maps high-dimensional atomistic coordinates r to lower-dimensional coordinates x through a coordinate transformation, here assumed linear: x = Ξr.The coarse-graining matrix Ξ clusters atoms into coarse-grained beads.
  • Coarse-grained representation: The learned energy function U(x; θ) is used with a dynamical model, with θ representing neural-network weights rather than classical potential parameters.The model can be simulated using dynamics such as Langevin dynamics.
  • Force matching: Force matching trains the model to predict instantaneous projected atomistic forces with minimal mean square error, achieving thermodynamic consistency under the employed dynamics.The model force is −∇U(x; θ), while the reference is the atomistic force projected onto coarse coordinates.
  • Statistical learning formulation: The force-matching error decomposes into noise, bias, and variance, where noise is fixed by the coarse-graining map and bias-variance balance governs model selection.Cross-validation estimates estimator error using repeated training-validation splits, while the PMF error measures mismatch between mean forces and predicted forces.
  • CGnets: CGnets learn a free energy and obtain forces through its gradient, enforcing a conservative potential of mean force.This architecture preserves the force-field structure required for a free energy-generated mean force.
  • Regularized CGnets: Regularized CGnets add a baseline energy whose constraint terms drive U toward infinity for unphysical states outside the training data.This prior energy is intended to prevent simulations from entering forbidden regimes with overstretched bonds or overlapping atoms.

2-dimensional toy model

The toy model tests force matching and cross-validation for learning a one-dimensional free energy from a rugged two-dimensional potential. The broader demonstrations show that regularized CGnets reproduce atomistic free-energy features better than few-body spline models, including multi-state protein behavior.

  • Toy model setup: The two-dimensional toy system uses a double-well potential along x, harmonic confinement along y, and small-scale ripples.
  • Toy model setup: The coarse-graining map projects the two-dimensional system onto x, whose exact coarse-grained free energy can be computed.
  • Learning and evaluation: A one-million-step trajectory trains feature regression and regularized or unregularized CGnets on x-direction forces, with cross-validation selecting hyperparameters.
  • Toy model results: In well-sampled x regions, feature regression and CGnet predictors accurately fit the exact mean force, while sparse boundaries produce divergence, especially for unregularized CGnets.A suitable prior energy makes the free energy increase outside the training region without affecting accuracy within it.
  • Toy model results: Free energies from learned-model simulations agree with the reference in the x range having significant equilibrium probability.The simulated free energies are obtained by histogramming trajectories generated with the learned models.
  • All-atom applications: For alanine dipeptide, regularized CGnets have lower cross-validation and free-energy-surface errors than spline models, while regularization avoids unphysical structures.The CGnet uses five central backbone atoms and no solvent; its multi-body architecture captures interactions absent from the spline model.
  • All-atom applications: For Chignolin, CGnet reproduces folded, unfolded, and partially misfolded minima, whereas the spline model cannot reproduce the folding/unfolding behavior.The reported correspondence is between CGnet minima and those observed in the all-atom solvated model.

Conclusions

The paper frames force matching as supervised learning and introduces CGnets that encode physical invariances, constraints, and prior energies. CGnets reproduce equilibrium distributions and complex solvation free energies, while larger systems, transferability, and dynamical reproduction remain limited.

  • Force matching becomes a supervised learning problem whose error decomposes into physically meaningful Bias, Variance, and Noise terms.Cross-validation can select hyper-parameters and rank coarse-graining models.
  • CGnets encode translational and rotational invariance, force equivariance, conservative mean forces, and prior energies that prevent unphysical simulations.Prior energies prevent divergence into states with overstretched bonds or clashing atoms.
  • CGnets automatically include multi-body effects and nonlinearities, providing better approximations than commonly used coarse-grained functional forms.The models can encode Chignolin’s complex solvation free energy using only protein Cα atoms.
  • The study is a proof of principle on relatively small solutes, and larger, more complex molecules may require additional physical-constraint terms.The authors identify extension to larger systems as an additional challenge.
  • The molecule-specific coarse-grained models are not transferable across different systems, leaving transferability–accuracy trade-offs for future work.The authors suggest particle-type and environment-dependent features as a possible route toward transferability.
  • The formulation targets structural properties but does not determine dynamical evolution equations, so reproducing fine-grained slow dynamics requires alternative approaches.

Supporting Information Available

The Supporting Information provides derivations, cross-validation results, training details, model distributions, and additional alanine dipeptide and Chignolin analyses.

  • The Supporting Information includes a derivation of equation (6), cross-validation for all models, and additional model-training details.
  • It reports bond and angle distributions for the different alanine dipeptide models and free-energy changes across hyper-parameters.
  • It provides energy decomposition for the alanine dipeptide CGnet and details of the Chignolin setup and simulation.
  • It includes Markov State Model analysis of Chignolin all-atom simulations.

Supplementary Material

The supplementary derivations explain how force-matching prediction error is decomposed by centering force estimates around the mean force and showing the mixed term vanishes.

  • The force-matching error decomposition is obtained by adding and subtracting the mean force and splitting the squared norm.
  • The mixed term is zero, making the resulting expression equivalent to Eq. (6).
  • The expected prediction-error decomposition similarly adds and subtracts the mean estimator f̄(X) = E[−∇U(X; θ)].
  • The remaining terms define bias and variance, while the element-wise product notation is used for the cross term.

Cross-validation for the coarse-graining of the 2d toy model

For the two-dimensional toy model, cross-validation selects features and CGnet hyper-parameters by minimizing cross-validation error, with uncertainty estimated across five folds.

  • Feature selection: Cross-validation selects the first four of twenty basis functions for feature regression in the two-dimensional toy model.The minimum cross-validation error is obtained when the first four functions are used.
  • Feature set: Table S1 contains twenty elementary basis functions used in the toy-model feature set.
  • Results: The CGnet toy-model cross-validation results are reported in Tables S2 and Figure S1.
  • Results: Table S2 optimizes the unregularized CGnet over network depth D and width W, reporting cross-validation error in (kBT)^2.The length unit is one.
  • CGnet hyper-parameters: Figure S1 reports feature selection and two-stage CGnet hyper-parameter selection for network depth D and width W.Red dashed lines indicate the minimum cross-validation error.

Training CG models

Networks were trained with Adam using PyTorch, with batch sizes tailored to the 2D and alanine-dipeptide models. Training and validation errors were monitored, while hyperparameters were selected by cross-validation.

  • Networks were optimized using Adam adaptive stochastic gradient descent with default settings in PyTorch.The batch size was 128 for the 2D model and 512 for alanine dipeptide.
  • Training and validation errors were evaluated for both the 2D toy model and alanine dipeptide.Errors were averaged over 200000 points; validation used a fixed validation set.
  • Hyperparameter choices were made through cross-validation rather than from training error alone.

Distribution of bond distances and angles for the different models of alanine dipeptide

For alanine dipeptide, regularized CGnet and spline models reproduce the all-atom distributions, whereas the unregularized CGnet produces a broader distribution. Free-energy profiles were compared across selected hyperparameter combinations.

  • Regularized CGnet and regularized spline models agree with the true all-atom distributions for the reported angles and bonds.The unregularized CGnet distribution spans a wider range.
  • Five CGnet hyperparameter combinations, C1 through C5, were used to examine how free-energy profiles approach the optimum.The C5 combination corresponds to the global optimum.
  • Four spline hyperparameter combinations, S1 through S4, were compared through their two-dimensional free-energy profiles.S4 corresponds to the global optimum.
  • Cross-validation error was compared with mean square free-energy difference for the selected CGnet and spline hyperparameters.

Energy decomposition for the CGnet model of alanine dipeptide.

The CGnet energy is decomposed into neural-network and baseline contributions, with both pointwise and dihedral-angle-binned averages reported. The Chignolin training data came from atomistic simulations and force calculations.

  • Energy decomposition: The total CGnet energy is decomposed into dense-network and baseline, or prior, energy contributions.The decomposition is reported for each simulated point and as averages over bins in dihedral-angle space.
  • Energy decomposition: The baseline energy enforces physical constraints and is treated as an important component of the CGnet model.
  • Atomistic simulation data: Chignolin simulations used the CHARMM22* force field, TIP3P water, a 350K temperature, and a 4 fs integration timestep.
  • Atomistic simulation data: Production sampling comprised 3744 independent 50 ns simulations, totaling 187.2 µs of aggregate simulation time.The simulations were run using the GPUGRID distributed computing platform.
  • Atomistic simulation data: Force data for CGnet training was computed for all atoms from the MD trajectories using the same parameters as the simulations.

Markov State Model analysis of Chignolin all-atom simulations

Chignolin atomistic trajectories were reduced to kinetic coordinates, clustered into discrete states, and used to estimate a Markov state model. The selected lag time was supported by approximately constant implied timescales above 20 ns.

  • Feature construction and dimensionality reduction: The trajectories were featurized into 45 non-nearest-neighbor Cα distances and reduced with TICA at a lag of τ = 25 ns.
  • Feature construction and dimensionality reduction: Four TICs retaining 95% kinetic variance were clustered into 350 discrete states using k-means.All MD data were mapped onto these states for Markov state model estimation.
  • MSM validation: MSM implied timescales became constant within statistical uncertainty for lag times above approximately 20 ns.
  • MSM validation: The Chignolin Markov state model was estimated at τ = 37.5 ns for convergence validation.The validation examined stationary probabilities of metastable states and MSM implied timescales.
  • CG-model hyperparameter selection: Five-stage cross-validation selected CG-model hyperparameters including network depth, width, excluded-volume parameters, and Lipschitz regularization strength.The selected values were used for the results reported in Fig. 7.
Loading 1812.01736v3…