Source-linked AI summary
Stochastic Control Policies for Robust Molecular Transition Path Sampling
Jingqian Liu, Yu-Hsiang Wang, Yanru Qu, Ge Liu
TL;DR
Rollout-based control methods can generate physically plausible molecular transition paths but show unstable, seed-dependent performance. This paper introduces force-space and latent-space stochastic policies, finding improved path quality and robustness across three biomolecular systems, with LaS-TPS strongest overall.
Problem
MD-rollout-based transition path control offers physically plausible trajectories but remains unstable and strongly sensitive to random initialization.
Method
The paper introduces FS-TPS, which samples state-dependent Gaussian force controls, and LaS-TPS, which decodes compact latent samples into correlated atom-wise force variations.
Results
Across alanine dipeptide, chignolin, and BBL, both stochastic policies improve transition-path quality and initialization robustness over deterministic control, with LaS-TPS strongest overall.
Takeaways & Limitations
Stochasticity placement is a consequential design choice for robust MD-rollout TPS, with compact latent control providing the strongest overall balance of transition success and energetic quality.
Takeaways & Limitations
MD-rollout methods trade off transition probability against energetic path quality and do not guarantee endpoint satisfaction.
Abstract
from arXiv · showhide
Transition path sampling (TPS) aims to efficiently generate rare molecular transition trajectories between metastable states and is essential for understanding biomolecular mechanisms. Beyond traditional molecular dynamics (MD)-based sampling, machine learning has become central to state-of-the-art TPS. One major class of methods learns control forces during explicit MD rollouts. By preserving the underlying molecular dynamics, these methods tend to produce more physically plausible trajectories than endpoint-conditioned generators that construct paths directly. However, rollout-based control methods have been reported to exhibit unstable and strongly seed-dependent performance. We recast rollout-based control as learning a path-space proposal distribution and investigate stochasticity placement as a design choice for improving exploration and optimization robustness. We develop two stochastic policies: FS-TPS, which directly parameterizes a state-dependent Gaussian distribution over the control policy output, and LaS-TPS, which samples a compact latent control variable and decodes it into structured, cross-atom-correlated force variation. We conduct extensive multi-seed experiments on three biomolecular systems of increasing size: alanine dipeptide, chignolin, and BBL, a fast-folding protein. Stochastic policies consistently improve transition success and path quality over deterministic-policy baselines while substantially reducing sensitivity to random initialization.
Introduction
The introduction frames TPS as a difficult rare-event sampling problem and recasts learning-guided MD rollouts as stochastic path-space proposal learning. It introduces FS-TPS and LaS-TPS to improve exploration, transition success, energetic performance, and robustness over random seeds.
- Motivation: TPS generates reactive trajectories between metastable states to provide mechanistic insight, but relevant free-energy barriers are rarely crossed.This makes efficient transition-path sampling challenging for biomolecular processes such as protein folding, conformational switching, and chemical reactions.
- Existing approaches: Learning-guided MD-rollout methods predict configuration-dependent bias forces that are combined with physical forces before advancing trajectories through an MD integrator.This preserves the molecular-dynamics rollout structure rather than constructing paths directly.
- Stochastic control: The work introduces stochasticity into policy learning to improve exploration and reduce instability and sensitivity to early replay-buffer trajectories.This perspective is motivated by action-space perturbations, parameter-space noise, entropy-regularized stochastic policies, and stochastic optimal control.
- FS-TPS: FS-TPS replaces the deterministic bias predictor with a state-dependent Gaussian policy that samples the applied control during each MD rollout.The introduction reports that FS-TPS improves RMSD and transition performance on alanine dipeptide, although the supplied passage truncates the quantitative results.
- LaS-TPS: LaS-TPS introduces stochasticity through a compact latent representation whose shared decoder produces correlated variations in full-atom forces.Under local linearization, the induced force covariance is correlated and state-dependent, with rank upper-bounded by the latent dimension.
- Results: Across three molecular systems, LaS-TPS achieves the strongest overall transition-success and energetic performance while maintaining robustness over random seeds.The systems are alanine dipeptide, chignolin, and BBL, a fast-folding protein.
Related work
Related work spans Monte Carlo path sampling with learned proposals, rollout-based learned control during explicit molecular-dynamics simulations, and endpoint-conditioned or path-space generative models that construct trajectories directly. These approaches replace or augment traditional trajectory-space moves through learned proposals, biasing forces, or generative path distributions.
- Monte Carlo path sampling and learned proposals: Monte Carlo transition path sampling uses shooting and shifting moves to generate reactive trajectories from existing ones, with efficiency depending critically on shooting-point quality.This dependence motivated learned proposal distributions that replace hand-designed moves.
- Rollout-based learned control: Rollout-based methods learn biasing forces that steer explicit molecular-dynamics trajectories toward a target basin.The reactive path ensemble was formulated as a variational control problem with a time-dependent force reweighting the unbiased path measure, while PIPS introduced collective-variable-free policy learning.
- Endpoint-conditioned and path-space generative models: Endpoint-conditioned models construct entire trajectories at once, enforcing both metastable endpoints by construction rather than reaching the target basin through simulation.One line formulates transition path sampling through Doob’s h-transform and learns endpoint-conditioned path distributions with simulation-free variational objectives.
- Endpoint-conditioned and path-space generative models: Other path-space generative methods use sequence-to-sequence, diffusion, flow, normalizing-flow, or trajectory-data-based models to generate transition paths.The cited approaches include fixed-window attention for longer paths, Onsager–Machlup action minimization, path normalizing flows, and direct trajectory-data training.
Methodology
The methodology formulates TPS control as stochastic path-distribution learning, with FS-TPS sampling state-dependent Gaussian forces and LaS-TPS sampling latent variables decoded into structured forces. Training combines path-measure matching with entropy or information-bottleneck regularization, while replay-buffer sampling and covariance analyses support implementation and characterization.
- TPS-DPS: TPS-DPS minimizes a log-variance objective between biased-dynamics and target transition-path distributions, enabling off-policy replay-buffer training with a learnable control variate.The control variate estimates the unknown normalization constant and reduces gradient-estimator variance.
- FS-TPS: FS-TPS models admissible bias forces as a conditional diagonal Gaussian and samples them with the reparameterization trick.The sampled force is u_t = µ_θ(x_t, x_B) + σ_θ(x_t, x_B) ⊙ ϵ, with ϵ ∼ N(0, I).
- FS-TPS: FS-TPS combines path-measure matching with conditional differential-entropy regularization to approach the target path measure while preventing premature collapse of force-distribution variability.The entropy coefficient is temperature-dependent through a tunable λ_ref parameter.
- LaS-TPS: LaS-TPS uses a variational information bottleneck to sample a compact diagonal-Gaussian latent control variable and decode it into the bias force.Its latent distribution is regularized toward a standard Gaussian prior with strength controlled by β_KL.
- LaS-TPS: A shared decoder maps low-dimensional latent perturbations to full atom-wise forces, producing generally non-diagonal, state-dependent covariance and correlated multi-atom control patterns.The stochastic force variation is restricted to the column space of the decoder Jacobian; conditional fluctuations are empirically sampled with K = 1024 queries across fixed transition-state banks.
Experiments
Experiments on alanine dipeptide, chignolin, and BBL show that stochastic rollout policies improve transition-path performance and robustness, with LaS-TPS delivering the strongest overall results among MD-rollout methods. Ablations indicate that entropy and KL regularization, along with structured latent stochasticity, are central to these gains.
- Experimental setup: Experiments evaluate TPS-DPS, FS-TPS, and LaS-TPS on alanine dipeptide, chignolin, and BBL using on-the-fly biased-MD path generation.The systems contain 22, 166, and 711 atoms, respectively, and represent increasing molecular size.
- FS-TPS performance: 69.79 ± 8.84% THP and 0.26 ± 0.05 ˚A RMSD are achieved by FS-TPS on alanine dipeptide, improving over TPS-DPS values of 45.67 ± 28.26% and 0.43 ± 0.22 ˚A.These results are averaged across nine random seeds, and consistent THP improvements are also observed on chignolin and BBL.
- FS-TPS ablations: Entropy regularization maintains stochastic exploration during replay-buffer construction, whereas removing it largely degrades FS-TPS performance.Scaling FS-TPS sampling noise from 0 to 1.5 at inference changes RMSD and THP little; only scale 10 substantially degrades ETS.
- FS-TPS ablations: Temperature-scaled Gaussian noise added to a deterministic policy fails to reproduce FS-TPS THP, indicating that learned state-dependent force distributions—not randomness alone—drive the improvement.The learned distribution is sustained by entropy regularization.
- Latent stochasticity: LaS-TPS produces low-effective-rank, cross-atom-correlated force fluctuations while maintaining broadly utilized latent dimensions, rather than collapsing its latent distribution.FS-TPS effective ranks are 52.56, 30.23, and 139.09 for alanine dipeptide, chignolin, and BBL, whereas LaS-TPS fluctuations are nearly one-dimensional.
- LaS-TPS performance: LaS-TPS achieves the strongest overall performance among MD-rollout methods, obtaining the lowest ETS on all three systems while retaining competitive RMSD and transition success.Its gains are not explained by greater model capacity, since it uses 20.8%, 13.4%, and 71.5% of corresponding baseline MLP parameter counts for alanine dipeptide, chignolin, and BBL.
Conclusion and Limitations
The paper introduces FS-TPS and LaS-TPS as stochastic-control formulations for MD-rollout-based TPS, finding improved robustness and transition-path quality relative to deterministic control across three molecular systems. It also identifies a trade-off between transition success and path quality, reflecting limitations of sequential MD rollouts and differences among TPS formulations.
- Stochastic-control formulations: FS-TPS models a state-dependent force-space distribution, while LaS-TPS samples a compact latent representation decoded into full atom-wise bias forces.Both methods inject stochasticity into MD-rollout-based transition path sampling, but at different stages of control generation.
- Empirical findings: Across three molecular systems, both stochastic approaches improve robustness to random initialization and transition-path quality relative to deterministic control.LaS-TPS provides the strongest overall balance of transition success and path quality among the approaches described.
- Trade-offs and limitations: LaS-TPS exhibits a trade-off between THP and ETS, with the balance changing as the KL coefficient varies.The passage presents this as a broader trade-off across TPS formulations rather than a uniquely LaS-TPS-specific behavior.
- Trade-offs and limitations: MD-rollout-based methods must discover the target basin through sequential simulation and therefore do not guarantee endpoint satisfaction.Endpoint-conditioned generators achieve high transition success by construction, but may produce paths with less favorable energy characteristics.
Appendix · Proof of the Structured-Covariance Proposition
The proof models the control output as a decoder applied to a state-conditioned Gaussian latent variable. A first-order expansion shows that latent variability induces structured, state-dependent output covariance within the decoder Jacobian’s column space.
- Proof of the Structured-Covariance Proposition: The control output is parameterized as u = gθ(z), with z | x, xB ∼ N(µz(x, xB), Σz(x, xB)).
- Proof of the Structured-Covariance Proposition: A first-order Taylor expansion of gθ around the state-dependent latent mean µz(x, xB) provides the proof’s local approximation.
- Proof of the Structured-Covariance Proposition: The latent residual has conditional mean E[z − µz(x, xB) | x, xB] = 0 and covariance Cov(z | x, xB) = Σz(x, xB).
- Proof of the Structured-Covariance Proposition: The resulting conditional mean and covariance of the decoded output are obtained approximately from the local expansion.
- Proof of the Structured-Covariance Proposition: The centered output satisfies u − E[u | x, xB] ≈ Jz(x, xB)(z − µz(x, xB)) and therefore lies in Col(Jz(x, xB)).
- Proof of the Structured-Covariance Proposition: Although Σz(x, xB) is diagonal, the induced output covariance can contain off-diagonal entries.
- Proof of the Structured-Covariance Proposition: These covariance entries are generally nonzero for i ≠ j, while state-dependent latent means and covariances make the output covariance and low-dimensional subspace vary with (x, xB).
Algorithm and implementation details of FS-TPS and LaS-TPS
FS-TPS and LaS-TPS share a common policy input and are trained and evaluated with explicit stochastic-policy procedures. Their implementation specifies sampled-policy scaling, method-specific mean policies, system-dependent hyperparameters, and replay-buffer sampling for LaS-TPS.
- Algorithms: Training and inference for FS-TPS and LaS-TPS are specified in Algorithms 1 and 2.The supplied implementation description identifies these algorithms as the procedures for both methods.
- Policy input: Both methods use the target conformation re-expressed in the current configuration’s reference frame as part of their policy input.The passages define the transformed target conformation and identify it as a common policy input.
- Policy evaluation: The default sampled-policy evaluation uses α = 1, whereas α = 0 yields the conditional mean-force policy for FS-TPS and the decoded-latent-mean policy for LaS-TPS.The α parameter controls evaluation of the stochastic policies and distinguishes the two deterministic mean-policy interpretations.
- Hyperparameters: FS-TPS uses system-specific βH values, while LaS-TPS evaluates βKL grids that vary across alanine dipeptide, chignolin, and BBL protein.The reported values are βH = 1 × 10−5, 5 × 10−6, and 5 × 10−5 for FS-TPS; βKL ∈{10−2, 10−3, 10−4} or βKL ∈{10−6, 10−7} for LaS-TPS.
- Replay buffer sampling: LaS-TPS minibatches combine trajectories from a low-RMSD pool with uniform samples from the full replay buffer.The low-RMSD pool is built from each trajectory’s closest target-state approach after Kabsch alignment, using a relaxed Gaussian indicator with bandwidth σ.
Evaluation Metrics
Evaluation uses ensembles of 64 sampled paths and combines structural accuracy, target-hit success, transition-state energy, and alanine-dipeptide mode coverage. These metrics are computed from re-evaluated trajectories and system-specific criteria.
- Evaluation protocol: All metrics are computed over an ensemble of 64 sampled transition paths, with forces and potential energies re-evaluated at every trajectory configuration.Atomic positions are recorded at each step before force-field re-evaluation.
- Structural accuracy: RMSD measures heavy-atom deviation between each path’s final configuration and the target basin reference after Kabsch optimal rigid-body superposition.This quantifies final structural accuracy relative to the target structure.
- Path success: THP is the percentage of paths whose final configurations satisfy system-specific target-hit criteria in low-dimensional collective-variable spaces.Alanine dipeptide uses periodic (ϕ, ψ) distances, while chignolin and BBL use distances in the leading two TIC coordinates.
- Transition-state energy: ETS is the maximum potential energy attained along trajectories classified as successful.The metric reports energy at the transition state for successful paths.
- Mode coverage: Alanine-dipeptide mode coverage tests whether successful paths visit both known saddle-point mechanisms connecting the C5 and C7ax basins.A path must pass within 0.3 rad of each saddle point in dihedral space at some point along its trajectory.
Computation of stochastic-force structure metrics · Per-state stochastic-force metrics
The paper computes five stochastic-force structure metrics from per-state force ensembles, separating internal molecular deformation from rigid-body motion and validating latent-space utilization for LaS-TPS. Metrics are evaluated across fixed molecular states and summarized across trained seeds.
- Computation of stochastic-force structure metrics: Five metrics quantify internal rank, dominant-mode variance, excess cross-atom covariance, stochastic amplitude relative to mean force, and LaS-TPS latent-rank utilization.The same procedures are applied to alanine dipeptide, chignolin, and BBL, with system-specific atom counts, latent dimensions, and evaluated-state counts.
- Per-state stochastic-force metrics: Per-state metric values are averaged over the state bank for each seed, then reported as the mean and standard deviation across independently trained seeds.The target conformation xB remains fixed across the state bank and is omitted from the notation.
- Per-state stochastic-force metrics: At each fixed configuration, the policy is queried KMC = 1024 times to form an applied bias-force ensemble for FS-TPS and LaS-TPS.FS-TPS samples directly from a force-space Gaussian, whereas LaS-TPS samples a fresh latent variable and decodes it into force.
- Per-state stochastic-force metrics: Internal-force metrics remove three translational and three mass-weighted rotational modes, so reported statistics capture molecular deformation rather than rigid-body motion.Covariance eigenvalues are computed from squared singular values of the projected, centered sample matrix for numerical stability.
- Per-state stochastic-force metrics: Internal effective rank measures the continuous effective number of equally weighted stochastic-force directions using the exponential of spectral Shannon entropy.It equals q when exactly q nonzero eigenvalues have equal magnitude.
- Per-state stochastic-force metrics: Internal top-mode explained variance reports the percentage of internal stochastic-force variance captured by the dominant eigenmode.Values near 100% indicate concentration along one internal direction, while smaller values indicate a broader spectrum.
- Per-state stochastic-force metrics: Excess cross-atom correlation compares blockwise cross-atom covariance against a permutation null that preserves marginal covariances while destroying coupling.The null is generated by independently permuting sample indices for each atom across B = 50 repetitions; values near zero indicate no structure beyond finite-sample background.
- Per-state stochastic-force metrics: LaS-TPS latent effective-rank utilization tests whether low-rank force covariance reflects decoder structure rather than collapse of sampled latent noise.Values close to one indicate approximately full-rank latent sampling; thus low-rank force covariance with Rz ≈1 indicates nonlinear-decoder structure.
Global effective rank across states … LaS-TPS
The paper defines a cross-state effective-rank measure for LaS-TPS and reports ablations, endpoint-generator comparisons, and training and inference procedures for FS-TPS and LaS-TPS. These results and algorithms emphasize stochastic-force structure, seed sensitivity, and rollout implementation.
- Global effective rank across states: Global effective rank near one indicates a shared dominant stochastic-force direction across states, whereas larger values indicate state-dependent directional changes.The metric is computed separately for each trained seed and summarized by mean and standard deviation across seeds.
- Additional results and ablation studies: 43.6% mean THP was achieved by the temperature-scaled Gaussian-noise baseline across nine seeds, with THP ranging from 3.1% to 70.3%.The baseline used σ = 10−3 and Tref = 300 K and remained strongly seed-sensitive.
- Additional results and ablation studies: Prioritized replay produced THP of 72.58 ± 5.01% versus 71.30 ± 8.24% without prioritization, while RMSD remained 0.25 ± 0.03 ˚A versus 0.25 ± 0.04 ˚A.ETS was 21.63 ± 5.49 kJ/mol with prioritized replay versus 18.54 ± 7.06 kJ/mol without it, so the ablation did not improve ETS.
- Additional results and ablation studies: Reducing LaS-TPS latent dimension to one increases sensitivity to initialization for alanine dipeptide.Table 6 reports individual-seed results and mean ± standard deviation over four seeds, with one degraded seed highlighted.
- Results of Doob’s Lagrangian and FAS: Doob’s Lagrangian generated 64 paths that all collapsed onto a single nearly straight trajectory in alanine dipeptide’s (ϕ, ψ) space.The paper compares Doob’s Lagrangian and FAS with MD-rollout TPS methods, noting that endpoint-conditioned metrics are not directly comparable.
- 4 FS-TPS: FS-TPS training predicts a conditional force distribution, draws fresh Gaussian noise, samples the bias force, and updates the replay buffer.The training algorithm generates biased-MD rollouts using annealing and then applies one of the stochastic control policies.
- LaS-TPS: LaS-TPS training predicts and samples a conditional latent distribution, decodes the latent variable, and augments the common objective with a policy-specific regularizer.The procedure samples path batches and forms a common log-variance objective before applying the stochastic-control regularizer.