Source-linked AI summary

Reweighted Autoencoded Variational Bayes for Enhanced Sampling (RAVE)

Joao Marcelo Lamim Ribeiro, Pablo Bravo Collado, Yihang Wang, Pratyush Tiwary

arXiv:1802.03420v1physics.chem-phcond-mat.softcond-mat.stat-mechphysics.comp-ph

TL;DR

Molecular simulations can remain trapped by high energy barriers, motivating enhanced-sampling methods that identify useful reaction coordinates and accelerate exploration. RAVE alternates molecular dynamics with variational-autoencoder learning, uses distribution matching to select an interpretable coordinate and bias, and reweights later simulations. It produced converged thermodynamic estimates, including a hydrophobic ligand-cavity binding profile obtained with at least 20-fold less computer time than umbrella sampling or metadynamics.

  • Problem

    High energy barriers trap molecular simulations, creating a need for enhanced sampling that can explore complex landscapes and identify effective reaction coordinates.

  • Method

    RAVE iterates between molecular dynamics and variational-autoencoder learning, matching latent and trial-coordinate distributions with Kullback-Leibler divergence before reweighted biased simulations.

  • Results

    At least 20 times less computer time than umbrella sampling and metadynamics was used for the hydrophobic ligand-cavity system while reproducing its binding free-energy profile.

  • Takeaways & Limitations

    RAVE provides a physically interpretable reaction coordinate and probability distribution simultaneously without relying on additional enhanced-sampling methods.

Abstract

from arXiv · show

Here we propose the Reweighted Autoencoded Variational Bayes for Enhanced Sampling (RAVE) method, a new iterative scheme that uses the deep learning framework of variational autoencoders to enhance sampling in molecular simulations. RAVE involves iterations between molecular simulations and deep learning in order to produce an increasingly accurate probability distribution along a low-dimensional latent space that captures the key features of the molecular simulation trajectory. Using the Kullback-Leibler divergence between this latent space distribution and the distribution of various trial reaction coordinates sampled from the molecular simulation, RAVE determines an optimum, yet nonetheless physically interpretable, reaction coordinate and optimum probability distribution. Both then directly serve as the biasing protocol for a new biased simulation, which is once again fed into the deep learning module with appropriate weights accounting for the bias, the procedure continuing until estimates of desirable thermodynamic observables are converged. Unlike recent methods using deep learning for enhanced sampling purposes, RAVE stands out in that (a) it naturally produces a physically interpretable reaction coordinate, (b) is independent of existing enhanced sampling protocols to enhance the fluctuations along the latent space identified via deep learning, and (c) it provides the ability to easily filter out spurious solutions learned by the deep learning procedure. The usefulness and reliability of RAVE is demonstrated by applying it to model potentials of increasing complexity, including computation of the binding free energy profile for a hydrophobic ligand-substrate system in explicit water with dissociation time of more than three minutes, in computer time at least twenty times less than that needed for umbrella sampling or metadynamics.

I. INTRODUCTION

RAVE addresses the difficulty of sampling molecular systems trapped by high energy barriers by jointly learning a reaction coordinate and its distribution through iterative simulation and deep learning. It targets physically interpretable, efficient enhanced sampling across increasingly complex systems.

  • High energy barriers can trap brute-force molecular simulations in limited regions of configuration space for extended periods.
  • RAVE combines reaction-coordinate identification with sampling its distribution through repeated molecular simulations and variational-autoencoder analysis.
  • Kullback-Leibler divergence selects a physically interpretable reaction coordinate from trial coordinates by matching their distributions to the learned latent-space distribution.
  • The selected reaction coordinate and distribution become the biasing protocol for subsequent simulations, with reweighting correcting statistics from biased trajectories.
  • RAVE differs from recent deep-learning methods by being independent of existing enhanced-sampling protocols and by filtering spurious local-minimum solutions.
  • Across model systems with barriers between 5 kBT and 30 kBT, RAVE achieved near-ergodic sampling and converged free-energy profiles accurately and efficiently.

1. Overview

The variational autoencoder models molecular-dynamics trajectories by learning a low-dimensional latent representation and reconstructing the original data. Its variational objective jointly trains encoder and decoder networks while regularizing the latent distribution against a prior.

  • RAVE uses a variational autoencoder to model molecular-dynamics trajectories through a low-dimensional latent representation.
  • The paper restricts the latent variable representation to one dimension, while noting that generalization is straightforward and reserved for future work.
  • The training objective combines an expected likelihood term with a Kullback-Leibler divergence penalty between the recognition model and prior.
  • The encoder maps high-dimensional data into latent space, while the decoder maps latent representations back into the original data space.
  • Maximizing the variational lower bound simultaneously trains the encoder and decoder to capture the main features of the data.

2. Neural Network Architecture

The neural-network architecture uses multilayer encoder and decoder networks with three hidden 512-dimensional layers. Inputs differ by application, and training uses RMSprop with a specified learning rate and epoch schedule.

  • Network architecture selection remains largely dependent on trial and error despite proposed systematic optimization approaches.
  • The input trajectories contain 200,000 two-dimensional datapoints for the model potentials and approximately 6,000 three-dimensional datapoints for fullerene unbinding.
  • Encoder hidden layers transform each datapoint through three sequential 512-dimensional vectors before producing Gaussian mean and variance parameters.
  • Decoder hidden layers transform a sampled one-dimensional Gaussian latent variable through three sequential 512-dimensional vectors before reconstructing the original data space.
  • Training uses Keras and RMSprop with a learning rate of 0.005, normally for 100 epochs.
  • Later fullerene-unbinding rounds require longer training because biased simulations produce comparatively large weights.

B. Reweighted Autoencoded Variational Bayes for Enhanced Sampling (RAVE)

RAVE iteratively identifies a physically interpretable reaction coordinate and its distribution, then uses both to bias subsequent simulations while reweighting biased data.

  • RAVE screens trial reaction coordinates by minimizing the KL divergence between the VAE-learned latent distribution and each projected MD-data distribution.The distributions are normalized and discretized over matching one-dimensional bins.
  • The bias combines reaction-coordinate identification and bias construction rather than treating them as separate tasks.The paper notes this automation as an advantage relative to conventional workflows.
  • The selected reaction coordinate and probability distribution jointly determine the bias potential for the next molecular-dynamics simulation.The biased dynamics use V_MD = V0(R) + V_bias(χ(R)).
  • RAVE reconstructs unbiased statistics from biased simulations using importance-sampling weights before the next iteration.Reweighting is applied both to trial reaction-coordinate projections and to the VAE inputs.
  • The VAE reconstruction loss uses datapoint weights, making latent-space learning correspond to reweighted simulation statistics.The weighted loss is a mean squared reconstruction error over the MD dataset.

A. Model Two-State Potential

On the Szabo-Berezhkovskii two-state potential, RAVE progressed from a trapped unbiased trajectory to an approximately ergodic biased simulation by updating the reaction coordinate and bias distribution.

  • The short unbiased simulation remained in its initial well and oscillated near the starting minimum at kBT = 1.This reflects the difficulty of crossing the potential barrier in the low-temperature regime.
  • The first biased simulation explored previously unsampled regions but became trapped in the second well because the initial bias was single-peaked.The single-peak bias lowered the barrier only in the forward direction.
  • θ = 30° became the optimum reaction coordinate after the first biased iteration, producing a two-peaked bias and several rapid transitions in the next simulation.The selected coordinate agreed with the analytical result and calculations from other methods.

B. Model Three-State Potential

On a three-state model potential, RAVE overcame an 8 kBT barrier after one iteration and produced substantially more ergodic trajectories with rapidly converging basin free-energy differences after five iterations.

  • The three-state model potential was used to test RAVE beyond the two-state example.
  • θ = 85° was selected as the optimum reaction coordinate, with a single-peaked bias generated from the learned latent distribution.
  • One RAVE iteration produced a quick transition from the initial well across the 8 kBT barrier separating two wells.
  • After five RAVE iterations, the trajectory became significantly more ergodic and the free-energy difference between basins converged extremely quickly relative to unbiased MD.

C. Hydrophobic Ligand-Cavity System in Explicit Water

RAVE was tested on fullerene-ligand unbinding in explicit water, where it converged on an interpretable reaction coordinate and reproduced benchmark free-energy results with substantially less computer time.

  • System and challenge: The system tests whether RAVE can overcome a ∼30 kBT barrier associated with a 200-second residence time and reproduce the free-energy profile efficiently.The corresponding unbiased-MD timescale is estimated at more than 1,000,000 years with available supercomputing resources.
  • System and challenge: The three order parameters are z, the axial separation ρ, and the cavity solvation state w.Each 0.5 ns MD round records these variables for the RAVE procedure.
  • Reaction-coordinate convergence: The optimized reaction coordinate has the form c_z z + c_ρ ρ + c_w w, with coefficients representing order-parameter weights.This parameterization makes the learned coordinate physically interpretable.
  • Reaction-coordinate convergence: After about 10 RAVE iterations, the reaction-coordinate weights converged near a prior reference, with z dominant, ρ secondary, and w nearly absent.After 22 rounds, the bias induced unbinding in multiple independent short MD runs.
  • Free-energy results: RAVE free-energy profiles agreed with two-dimensional umbrella sampling and one-dimensional metadynamics across the binding profile.The comparison used profiles from two independent final RAVE rounds.
  • Free-energy results: At least 20 times less computer time was used by RAVE than reported for umbrella sampling and metadynamics.The example also extracts a physically relevant coordinate capturing steric and solvation effects.

D. General Comments on the Usage of RAVE

The authors describe a two-stage heuristic for rejecting misleading VAE solutions, motivated by local-minimum trapping and ambiguity among similar loss values.

  • Spurious-solution filtering: Deep-learning protocols can become trapped in local minima, producing misleading solutions with deceptively similar loss functions.This creates uncertainty about which learned solution is correct.
  • Spurious-solution filtering: RAVE first ranks candidate solutions by maximal bias while requiring zero bias in unsampled regions.For the bias shown in Fig. 4(b), the maximal bias is approximately 5 kBT; many spurious VAE solutions have lower values and are rejected.
  • Spurious-solution filtering: If multiple solutions pass the maximal-bias screen with similar values, the authors apply a second selection step.The supplied passage ends before specifying that second step.

IV. DISCUSSION

The discussion frames RAVE as a VAE-based method that prioritizes a useful latent-space distribution over an exact latent coordinate, while allowing explicit trade-offs between intuition and computational cost.

  • Method rationale: RAVE uses the VAE-learned latent-space probability distribution as its benchmark for evaluating candidate reaction-coordinate proxies.The authors do not seek to reproduce the full complicated nonlinear true reaction coordinate.
  • Method rationale: Physically interpretable reaction-coordinate proxies can approximate relevant features of a complicated true coordinate without reproducing all configurational dependence.This motivates replacing the precise latent variable with a more intuitive representation.
  • Method flexibility: RAVE can use more complex reaction coordinates to match the VAE probability distribution more accurately, trading intuition against computational cost.Using the latent variable directly would retain complicated nonlinear variables with less physical intuition.
  • Method flexibility: The maximal-bias heuristic is inspired by SGOOP: under constant diffusivity, deeper energy basins imply higher first-passage times and spectral gaps.The authors are exploring a possible connection between RAVE and SGOOP through more refined screening.
  • Scope and outlook: Across the applications, RAVE was much faster than unbiased MD and approximately 20 times faster than metadynamics or umbrella sampling for hydrophobic unbinding.The authors report converged thermodynamic observables with limited prior intuition and computational workload.
  • Scope and outlook: The demonstrated systems use only two or three order parameters, although initial tests suggest this number could be increased significantly.Future directions include adding temporal identity through time-lagged autoencoders.
Loading 1802.03420v1…