Source-linked AI summary

Coarse-Graining Auto-Encoders for Molecular Dynamics

Wujie Wang, Rafael Gómez-Bombarelli

arXiv:1812.02706v2physics.chem-phcs.LGstat.ML

TL;DR

Atomistic molecular dynamics can be computationally infeasible at the spatial and temporal scales relevant to materials, while coarse-graining requires both a reduced mapping and a coarse-grained Hamiltonian. Autograin uses an auto-encoder with reconstruction loss and force matching to learn these tasks jointly, and demonstrates the procedure on molecular and bulk-phase model systems.

  • Problem

    Atomistic simulations can be computationally infeasible at material-relevant scales, and coarse-graining mappings are generally hand-tuned despite their importance for consistent dynamics, structure, and thermodynamics.

  • Method

    Autograin jointly learns coarse-grained coordinates and their potential using an auto-encoder, reconstruction loss, and instantaneous force matching.

  • Results

    The learned mappings recover chemically interpretable groupings and show good agreement between coarse-grained and mapped atomistic equilibrium distributions and correlations across model systems.

  • Takeaways & Limitations

    Treating coarse-grained coordinates as latent variables enables a jointly trained encoding, deterministic decoding, and potential for simulations of larger systems over longer times.

  • Takeaways & Limitations

    Force-matching approaches are not guaranteed to capture non-equilibrium transport properties or transfer across different thermodynamic conditions.

Abstract

from arXiv · show

Molecular dynamics simulations provide theoretical insight into the microscopic behavior of materials in condensed phase and, as a predictive tool, enable computational design of new compounds. However, because of the large temporal and spatial scales involved in thermodynamic and kinetic phenomena in materials, atomistic simulations are often computationally unfeasible. Coarse-graining methods allow simulating larger systems, by reducing the dimensionality of the simulation, and propagating longer timesteps, by averaging out fast motions. Coarse-graining involves two coupled learning problems; defining the mapping from an all-atom to a reduced representation, and the parametrization of a Hamiltonian over coarse-grained coordinates. Multiple statistical mechanics approaches have addressed the latter, but the former is generally a hand-tuned process based on chemical intuition. Here we present Autograin, an optimization framework based on auto-encoders to learn both tasks simultaneously. Autograin is trained to learn the optimal mapping between all-atom and reduced representation, using the reconstruction loss to facilitate the learning of coarse-grained variables. In addition, a force-matching method is applied to variationally determine the coarse-grained potential energy function. This procedure is tested on a number of model systems including single-molecule and bulk-phase periodic simulations.

I. INTRODUCTION

Coarse-graining reduces molecular complexity and computational cost by averaging fast motions, but its mapping from atoms to coarse variables remains a critical challenge. The paper proposes learning this mapping from latent representations while jointly optimizing the coarse-grained potential.

  • Motivation: Coarse-grained models compress atomistic systems into fewer pseudo atoms, focusing on slow collective motions while averaging out fast local motions.This reduction enables lower-cost simulations of complex molecular processes.
  • Motivation: Existing coarse-graining approaches primarily parametrize potentials from atomistic simulations or experimental statistics, whereas the atom-to-coarse mapping is a separate central problem.The mapping is important for recovering consistent dynamics, structural correlations, and thermodynamics.
  • Approach: The paper frames coarse-grained variables as latent representations that can capture complex molecular data with fewer variables.This formulation draws on machine learning approaches for discovering hidden structure in complex datasets.
  • Approach: Autograin uses a constrained auto-encoder to compress atomistic molecular-dynamics data into a rigorously coarse-grained 3D representation.A reconstruction loss encourages the representation to preserve salient collective features.
  • Approach: The framework simultaneously learns the coarse-grained mapping and potential by combining reconstruction with instantaneous force matching.Force matching variationally identifies a potential reproducing the instantaneous mean force on the all-atom training data.

II. RESULTS

Autograin jointly learns coarse-grained mappings and force fields from atomistic trajectories, recovering chemically meaningful representations and structural correlations across molecules and liquids. Simple potentials show reasonable agreement, while neural potentials reproduce structural correlations almost exactly, subject to information loss and force-matching limitations.

  • Framework: Autograin jointly trains an auto-encoder and force-matching task to learn coarse-grained mappings and force fields from atomistic trajectories.The framework reconstructs all-atom data through a low-dimensional bottleneck while fitting instantaneous mean forces.
  • Single-molecule systems: OTP maps into three beads by grouping each phenyl ring, while aniline maps into two beads partitioning its amino and carbon environments.These mappings emerge from optimization rather than being manually specified.
  • Single-molecule systems: The learned OTP and aniline mappings reproduce held-out bond and angle distributions with good agreement between coarse-grained and mapped atomistic trajectories.For OTP, agreement is reported for each degree of freedom; aniline shows good bond-distribution agreement.
  • Bulk liquids: Neural potentials reproduce structural correlation functions almost exactly by representing complex correlations that simple classical functional forms cannot capture.The learned neural potential combines harmonic-like short-range behavior with non-bonded longer-range behavior.
  • Limitations: Deterministic decoding cannot reconstruct lost rotational degrees of freedom, producing averaged structures rather than reference instantaneous configurations.Force matching also does not guarantee perfect recovery of individual pair correlations, and simple potentials may lack higher-order flexibility.
  • Bulk liquids: Methane coarse-grained into one pseudo-atom achieves nearly perfect pair-correlation agreement, whereas ethane coarse-grained into two beads shows reasonable agreement with a classical force field.The ethane model uses same-type pseudo-atoms and classical bonded and non-bonded terms.

III. DISCUSSION

The discussion identifies limitations of deterministic decoding, force matching, fixed topologies, and thermodynamic transferability, while outlining extensions and broader simulation applications.

  • Limitations and future directions: Deterministic encoding and decoding irreversibly lose information, yielding average reconstructed structures rather than reference instantaneous configurations.A probabilistic auto-encoder with predictive atomistic back-mapping is proposed as a future extension.
  • Limitations and future directions: Force matching does not guarantee recovery of individual pair correlations because simple potentials may lack complex terms and structural cross-correlation effects.Iterative force matching and relative entropy are suggested to incorporate structural cross-correlations.
  • Limitations and future directions: The current model requires a predetermined topology to calculate total potential energy, limiting automatic learning of multi-particle force fields.Future work could probabilistically generate force-field topologies while optimizing the coarse-graining encoding.
  • Limitations and future directions: Bottom-up force-matching methods are not guaranteed to capture non-equilibrium transport properties or transfer across thermodynamic conditions.The proposed data-driven framework is described as enabling learning across different thermodynamic conditions and time-series training for transport properties.
  • Broader implications: Autograin jointly trains latent coarse-grained coordinates, deterministic decoding, and a transferable potential for larger systems and longer simulations.The framework is presented as a statistical-learning bridge across multiple simulation scales.

IV. METHODS

Autograin uses semi-supervised auto-encoder training to learn both an all-atom-to-coarse-grained mapping and a coarse-grained potential for new simulations.

  • Architecture and training: Autograin combines unsupervised reconstruction with supervised force matching to shape latent coarse-grained variables.The learned potential is defined in coarse-grained coordinates for later simulation.
  • Architecture and training: The framework learns an all-atom-to-coarse-grained mapping function and a potential usable for larger systems at lower computational cost.The approach is described as semi-supervised and variational.

A. Coarse-Graining Auto-encoding

Autograin learns physically meaningful coarse-grained coordinates by compressing atomistic Cartesian configurations through a constrained auto-encoder. Its encoder uses discrete assignment and its decoder reconstructs the original space, while deterministic compression necessarily loses information.

  • Encoding architecture: Autograin constrains latent variables to retain molecular structural information while representing coarse-grained coordinates with physical meaning.The latent space is designed to represent molecular positions and momenta rather than arbitrary abstract features.
  • Encoding architecture: The encoder maps atomistic coordinates in R3n to N coarse-grained particles in R3 using a linear Cartesian transformation.Here n is the number of atoms and N is the desired number of coarse-grained particles.
  • Discrete assignments: Gumbel-Softmax masks encoder weights so each atom asymptotically contributes to at most one coarse-grained variable as inverse temperature increases.The tunable fictitious inverse temperature β is gradually increased during training to learn discrete assignments.
  • Decoding and loss: The decoder uses an n-by-N matrix to map coarse-grained variables back into the original atomistic space.This provides a simple deterministic decoding approach for reconstructing atomistic configurations.
  • Decoding and loss: Deterministic low-dimensional reconstruction creates an irreversible information bottleneck, so decoded structures may represent averages rather than instantaneous atomistic configurations.The paper notes that this limits recovery of structural degrees of freedom removed by coarse-graining.

B. Variational Force Matching

Autograin jointly learns a coarse-grained mapping and potential by replacing constrained-dynamics mean-force estimation with instantaneous force matching. The resulting variational objective optimizes both the potential and encoder from atomistic data.

  • Variational formulation: The method conditions an instantaneous force-matching functional on the encoder to learn VCG over simultaneously learned coarse-grained coordinates.This couples optimization of the coarse-grained potential with the representation E(x).
  • Mean-force construction: The coarse-grained distribution and potential of mean force are defined over variables z obtained from the atomistic-to-coarse-grained mapping.The formulation introduces p(z) and A(z) as the coarse-grained distribution and corresponding many-body potential of mean force.
  • Mean-force construction: Instantaneous forces Finst(z) are conditional microscopic-force observables whose conditional expectation equals the mean force F(z), although their individual values are not unique.The specific instantaneous force depends on the choice of b, while conditional averages recover the same mean force.
  • Optimization target: The force-matching objective minimizes mean-square error between the mean force and the negative gradient of VCG.The potential parameters θ determine the coarse-grained force through automatic differentiation of VCG.
  • Optimization target: Minimizing Linst jointly over VCG(z) and E(x) yields a variational procedure for learning both the coarse-grained mapping and its force field.The paper relates Linst to the conventional objective L plus an additional instantaneous-force error term.

C. Model Training

Model training combines reconstruction and instantaneous force-matching losses, first learning a representative mapping and then jointly refining the encoder, decoder, and coarse-grained potential. The study compares classical and neural potential forms with different speed, flexibility, and transferability trade-offs.

  • Joint optimization: The total objective is the joint loss LV CGE = LAE + Linst, combining reconstruction and instantaneous force matching.The optimization stack uses both losses to train the coarse-grained representation and potential.
  • Joint optimization: Training uses atomistic trajectories with per-atom forces, and parameters are optimized by feed-forward propagation and backpropagation.The reconstruction and force-matching losses are minimized together during supervised training.
  • Training workflow: The workflow first pretrains the auto-encoder unsupervised, then jointly trains force matching with the auto-encoder to refine E(x), D(z), and VCG.This staged procedure produces a final mapping together with associated force-field parameters.
  • Potential choices: Classical Lennard-Jones and harmonic bonded potentials are fast to evaluate and transferable, but may lack the expressiveness needed for complex many-body potentials of mean force.These forms are presented as useful for simple phenomenological models and virtual screening.
  • Potential choices: Neural-network potentials fit complex many-body potentials more flexibly and faithfully, at the cost of poor transferability.The paper contrasts their flexibility with the simpler classical functional forms.

D. Computational details

The computational study uses OPLS-based trajectories, short gas-phase single-molecule simulations, and liquid methane and ethane systems. Models are trained on smaller systems and validated on coarse-grained liquid systems at matched density.

  • Trajectory generation: OPLS trajectories provide the molecular data, with gas-phase single-molecule simulations using 6 ps and 3000 frames from Langevin dynamics.The reported friction coefficient is 1 ps−1, and models are pretrained for 100 epochs with mini-batches of size 10.
  • Single-molecule systems: For single-molecule OTP, force matching learns harmonic bond and angle parameters from atomistic forces.Structural distributions are evaluated through normalized Boltzmann probabilities for bond and angle coordinates.
  • Liquid systems: Liquid training trajectories contain 64 methane molecules at 100 K and 64 ethane molecules at 120 K under NVT conditions.The liquid models use a SchNet-based neural potential with two convolutions.
Loading 1812.02706v2…