Source-linked AI summary
Data-Driven Collective Variables for Enhanced Sampling
Luigi Bonati, Valerio Rizzi, Michele Parrinello
TL;DR
Choosing collective variables from metastable-state information remains challenging, especially for nonlinear processes. The paper introduces Deep-LDA, a neural-network dimensionality-reduction method, and shows that it identifies variables that reproduce slow modes and promote transitions along chemically meaningful paths.
Problem
Collective variables must be obtained from limited metastable-state information while retaining discriminative power for enhanced sampling.
Method
Deep-LDA compresses metastable-state descriptors with nonlinear neural-network reduction followed by a linear transformation optimized by Fisher’s linear discriminant.
Results
Deep-LDA reproduces alanine dipeptide’s slow dihedrals and identifies aldol-reaction variables whose isolines correlate with the minimum-free-energy path.
Takeaways & Limitations
Feature analysis identifies the aldol reaction’s carbon-carbon bond and two proton-transfer contacts as the most relevant descriptors, providing chemical insight into the process.
Abstract
from arXiv · showhide
Designing an appropriate set of collective variables is crucial to the success of several enhanced sampling methods. Here we focus on how to obtain such variables from information limited to the metastable states. We characterize these states by a large set of descriptors and employ neural networks to compress this information in a lower-dimensional space, using Fisher's linear discriminant as an objective function to maximize the discriminative power of the network. We test this method on alanine dipeptide, using the non-linearly separable dataset composed by atomic distances. We then study an intermolecular aldol reaction characterized by a concerted mechanism. The resulting variables are able to promote sampling by drawing non-linear paths in the physical space connecting the fluctuations between metastable basins. Lastly, we interpret the behavior of the neural network by studying its relation to the physical variables. Through the identification of its most relevant features, we are able to gain chemical insight into the process.
Alanine dipeptide
For alanine dipeptide, Deep-LDA learns a low-dimensional collective variable from heavy-atom distances that reproduces the slow dihedral degrees of freedom and supports efficient enhanced sampling. Its behavior remains robust across network architectures, regularization choices, and sampling algorithms.
- Alanine dipeptide: Deep-LDA is trained from 45 heavy-atom distances collected in short unbiased trajectories from the two metastable basins.The resulting model can be loaded into PLUMED2 for enhanced-sampling simulations.
- Alanine dipeptide: Enhancing the Deep-LDA CV produces highly diffusive behavior comparable to directly biasing the Ramachandran angles φ and ψ.The free-energy surface is reported along the learned CV.
- Alanine dipeptide: Deep-LDA rapidly converges to the reference free-energy difference and reproduces φ and ψ as the slowest degrees of freedom.The comparison uses standard calculations based on φ and ψ, with the full two-dimensional landscape reported separately.
- Alanine dipeptide: The learned CV’s isolines closely reflect the free-energy surface despite lacking a direct correspondence with the individual dihedral angles.The isolines are computed from p(s | φ, ψ) using an OPES ensemble with a uniform target distribution in (φ, ψ).
- Alanine dipeptide: Similar results across neural-network architectures, regularization parameters, and Well-Tempered Metadynamics demonstrate robustness to these methodological choices.The finding indicates that performance is not tied to one parameterization or enhanced-sampling algorithm.
SUPPORTING INFORMATION
The supporting information details Deep-LDA implementation and tests its enhanced-sampling performance on alanine dipeptide and an aldol reaction. Deep-LDA reproduces relevant free-energy behavior, converges robustly, and drives transitions using chemical-distance inputs.
- Implementation: Deep-LDA uses regularized Fisher training with ReLU activations, ADAM optimization, and 10,000 configurations divided into batches of 2,000.The within-class scatter regularization uses λ = 0.05, while the weight L2 regularization uses γ = 10^-5 and the learning rate is 10^-4.
- Alanine dipeptide: Using only scalar distances, Deep-LDA constructs a nonlinear combination that distinguishes alanine dipeptide states and drives transitions similarly to the φ and ψ dihedrals.The comparison reports transition behavior against simulations biasing the two Ramachandran angles.
- Alanine dipeptide: Deep-LDA reproduces the topology of alanine dipeptide’s free-energy landscape in the Ramachandran plot from short unbiased simulations.The conditional distribution p(s | φ, ψ) was computed using OPES with a uniform target distribution in {φ,ψ} space.
- Alanine dipeptide: The reconstructed free-energy error is smaller than 0.5 kBT (1.25 kJ/mol) in the most relevant alanine dipeptide regions.The estimate is based on 10 independent simulations and comparison with a reference FES obtained by directly biasing φ and ψ.
- Convergence tests: All tested alanine dipeptide cases show quick convergence of the free-energy difference within the 0.5 kBT uncertainty across sampling methods and neural-network realizations.The convergence protocol uses 10 simulations per case, discards the first 2 ns, and updates ΔF every 1 ns.
- Aldol reaction: A Deep-LDA CV trained on all distances drives the aldol system between basins, although slightly less efficiently than a CV trained with contacts.The distance-based CV is transformed as s′ = s + s3 and used with OPES.