Source-linked AI summary

Bayesian ensemble refinement by replica simulations and reweighting

Gerhard Hummer, Jürgen Köfinger

arXiv:1509.04447v2physics.data-ancond-mat.stat-mechphysics.bio-phphysics.chem-ph

TL;DR

The paper addresses how to infer dynamic and partially disordered biomolecular ensembles from diverse measurements, including ensemble averages. It develops Bayesian, replica, reweighting, and adaptive formulations, showing their equivalence in appropriate limits and identifying linear replica-bias scaling for convergence to the optimal distribution.

  • Problem

    Ensemble refinement must characterize high-dimensional, dynamic biomolecular configuration spaces from limited heterogeneous measurements, including observables averaged over ensembles.

  • Method

    The paper formulates ensemble refinement as Bayesian posterior-functional optimization and analyzes EROS reweighting, replica simulations, BioEn, and adaptive sampling.

  • Results

    The methods are equivalent in appropriate limits, with replica and EROS convergence behaving as 1/N and 1/n, respectively, while replica bias must scale linearly with N.

  • Takeaways & Limitations

    Bayesian ensemble refinement provides a framework for integrating diverse experiments and sampling or reweighting distributions consistent with ensemble observables.

  • Takeaways & Limitations

    The procedure assumes a reference distribution, and estimating measurement and calculation errors can require uncertain error models or additional treatment.

Abstract

from arXiv · show

We describe different Bayesian ensemble refinement methods, examine their interrelation, and discuss their practical application. With ensemble refinement, the properties of dynamic and partially disordered (bio)molecular structures can be characterized by integrating a wide range of experimental data, including measurements of ensemble-averaged observables. We start from a Bayesian formulation in which the posterior is a functional that ranks different configuration space distributions. By maximizing this posterior, we derive an optimal Bayesian ensemble distribution. For discrete configurations, this optimal distribution is identical to that obtained by the maximum entropy "ensemble refinement of SAXS" (EROS) formulation. Bayesian replica ensemble refinement enhances the sampling of relevant configurations by imposing restraints on averages of observables in coupled replica molecular dynamics simulations. We show that the strength of the restraint should scale linearly with the number of replicas to ensure convergence to the optimal Bayesian result in the limit of infinitely many replicas. In the "Bayesian inference of ensembles" (BioEn) method, we combine the replica and EROS approaches to accelerate the convergence. An adaptive algorithm can be used to sample directly from the optimal ensemble, without replicas. We discuss the incorporation of single-molecule measurements and dynamic observables such as relaxation parameters. The theoretical analysis of different Bayesian ensemble refinement approaches provides a basis for practical applications and a starting point for further investigations.

I. INTRODUCTION

Ensemble refinement addresses the inverse problem of characterizing dynamic or partially disordered biomolecular structures from limited, heterogeneous experimental data. The Bayesian formulation preserves a reference ensemble while requiring appropriate ensemble-averaged observables to match measurements.

  • Ensemble refinement characterizes high-dimensional molecular configuration spaces from limited experimental information.
  • Experimental methods report diverse information, including molecular size, distances, local chemical environments, and global structure.
  • Bayesian priors regularize ill-conditioned, underdetermined inverse problems by encoding expectations about models and parameters before new data.
  • Single-copy refinement concentrates probability near configurations individually matching an observation, whereas ensemble refinement redistributes populations between reference rotamer states while matching the ensemble average.
  • The paper derives Bayesian ensemble distributions, connects discrete reweighting to EROS, and develops replica and adaptive practical approaches.

A. Bayesian single-copy refinement in configuration space

Single-copy Bayesian refinement ranks individual configurations by their consistency with measurements and a reference distribution. It models measured observables and uncertainties through a likelihood, then combines them with the prior to form the posterior.

  • Single-copy refinement assumes one configuration can explain all measured data and ranks configurations by their consistency with observations and prior.
  • Configurations x represent molecular states, while observables yi(x) are predicted for each state and compared with measured values y(obs)i.
  • The posterior combines the reference distribution p0(x) with the likelihood of observing the measured data given configuration x.
  • Gaussian likelihoods use measurement standard deviations, while correlated errors are represented through a covariance matrix Σ.
  • Posterior sampling can use an effective energy in equilibrium simulations or reweight configurations drawn from the reference distribution by their likelihood.

B. Bayesian ensemble refinement in probability density space

Bayesian ensemble refinement treats the probability density itself as the inferred object, combining a reference distribution with likelihoods based on ensemble-averaged observables. Variational optimization yields optimal distributions and discrete weights related to maximum-entropy reweighting.

  • The ensemble formulation treats p(x) as a probability density defining an ensemble, so prior, likelihood, and posterior become functionals of p(x).
  • A Kullback-Leibler divergence relative to p0(x) regularizes deviations from the reference ensemble, with θ controlling confidence in that reference.
  • Variational maximization produces an optimal Bayesian density, while normalization is enforced with a Lagrange multiplier.
  • For θ = 1, common-property refinement recovers the single-copy Bayesian posterior, and discrete reweighting converges to optimal Bayesian weights as n →∞.
  • Ensemble observables are averages over structures, and Gaussian-error likelihoods can be generalized to correlated errors and more general likelihood functions.
  • The optimal density generally satisfies a nonlinear integral equation that is difficult to solve in high-dimensional problems, motivating an adaptive replica-free algorithm.

C. Bayesian ensemble refinement in configuration space: The replica method

Replica refinement approximates ensemble distributions by coupling multiple simulated copies through restraints on averaged observables. Its large-replica limit reaches the optimal Bayesian distribution only with bias strength scaled linearly in the replica count.

  • Replica refinement simulates N copies in parallel and biases their averaged observables toward experimental measurements.
  • The method is designed to recover common-property refinement at N = 1 and optimal Bayesian ensemble refinement as N →∞.
  • Individual replicas converge to the optimal Bayesian configuration-space distribution in the infinite-replica limit.
  • Without the factor N, replicas decouple; related maximum-entropy analyses therefore differ in how restraint strength scales with replica number.
  • Linear scaling of the χ2 bias with N is the unique size-consistent choice: sub-linear scaling loses the bias, whereas super-linear scaling imposes effectively strict constraints.

D. BioEn method combining replica simulations with EROS

BioEn combines replica simulations with EROS to accelerate convergence toward the optimal Bayesian ensemble. It reweights replica states to correct finite-replica effects and supports optimal ensembles across θ values without running replica simulations for each one.

  • D. BioEn method combining replica simulations with EROS: BioEn combines EROS with replica simulations to accelerate convergence toward the optimal Bayesian ensemble.The combination also helps select θ values without running Bayesian replica simulations for all of them.
  • D. BioEn method combining replica simulations with EROS: Large-N replica sampling is computationally expensive, whereas EROS can use many structures efficiently but may converge slowly when reference and refined ensembles have little overlap.These complementary limitations motivate combining the two approaches.
  • D. BioEn method combining replica simulations with EROS: BioEn combines unbiased and biased N-replica simulations to obtain representative replica-state samples for reweighting.Unbiased simulations follow the product reference distribution, while biased coupled simulations use distributions pN(x1, . . . , xN).
  • D. BioEn method combining replica simulations with EROS: Relative replica-state weights are estimated across multiple simulations using iteratively determined free energies and biasing potentials.The biasing potential is Ui({xα}) ≡ Nχ2/2θi in biased runs and zero in unbiased runs.
  • D. BioEn method combining replica simulations with EROS: EROS reweighting corrects finite-replica effects in ensembles enriched by biased replica simulations.The resulting configurations and estimated weights are used as input for EROS refinement.

A. Illustrative examples of ensemble reweighting

Analytically tractable examples show how optimal Bayesian ensemble refinement relates to replica refinement and EROS, including convergence and the need for replica-number scaling. The BioEn combination reduces systematic errors at small replica counts.

  • Mean reweighting: The optimal Bayesian density remains Gaussian with reference variance s2, while its mean shifts toward the measured ensemble average Y according to s2 and σ2.The mean approaches Y when σ ≪ s and remains near the reference mean when σ ≫ s.
  • Replica refinement: As N → ∞, Bayesian replica refinement converges to the optimal Bayesian ensemble, with logarithmic density error scaling as O(1/N).For the Gaussian mean example, the asymptotic expansion is ln[qN(x)/p(opt)(x)] = f(x)/N + O(1/N^2).
  • Replica refinement: Scaling the χ2 restraint by N is necessary for size consistency; without it, the replica mean approaches the reference mean as N increases.The unscaled restraint yields a mean Y s2/(s2 + Nσ2), which tends to zero for σ > 0.
  • Second-moment reweighting: For the nonlinear second-moment observable, the optimal variance t2 approaches Y as σ → 0 and the reference variance s2 as σ → ∞.Intermediate uncertainty produces a nonlinear interpolation between these limits.
  • Double-well convergence: In the double-well model, replica reweighting shifts population from the left well to the right while largely preserving each well’s reference shape.For N = 2, the mean restraint instead pulls one replica into the barrier region.
  • EROS and BioEn: EROS converges to the optimal ensemble with logarithmic error scaling as 1/n, while BioEn removes substantial systematic replica errors for N = 2, 4, and 8.After EROS reweighting, the remaining errors are small and primarily statistical.

B. Practical considerations

The practical framework combines common-property and ensemble-average data, supports single-molecule and dynamic observables, and offers procedures for reweighting, parameter selection, and numerical solution.

  • Common-property and ensemble-average restraints can be combined by multiplying their likelihood terms in the posterior.
  • Data from single-molecule experiments: Single-molecule data can be incorporated by comparing experimental FRET-efficiency histograms with histogram observables calculated from an ensemble.The resulting χ2 can be used with both EROS and Bayesian replica ensemble refinement.
  • Dynamic, time-dependent data: Dynamic observables such as NMR NOEs and relaxation parameters can be treated approximately by generating trajectory segments around sampled configurations.The trajectory length τ should cover relevant correlation times so the observables can be estimated accurately.
  • Dynamic, time-dependent data: Time scaling can be optimized alongside ensemble refinement to account for incorrect simulation-model timescales, such as water-model viscosity errors.
  • Solving the EROS equations: The EROS weights can be obtained by numerical minimization or fixed-point iteration, with geometric mixing and θ sweeps improving stability and convergence.Starting at large θ and reducing it iteratively supports refinements across a broad confidence-factor range.
  • Choosing the confidence factor θ: θ expresses confidence in the reference distribution and can be selected using relative entropy versus χ2, energy-error expectations, or a nuisance-parameter treatment.EROS reweighting can estimate relative entropy efficiently across θ values, including when replicas generated the configurations.

IV. CONCLUSIONS

The paper develops and relates Bayesian approaches for refining molecular ensembles against diverse experimental data. It establishes optimal-distribution equivalences and convergence conditions while identifying practical assumptions and extensions.

  • The optimal Bayesian distribution for discrete configurations is identical to the distribution obtained with the EROS formulation.
  • Replica refinement requires scaling the biasing potential linearly with replica number N to remain size-consistent and converge to the optimal result as N →∞.
  • BioEn combines replica sampling, free-energy weighting, and EROS to address limited reference-to-optimal overlap and the use of relatively small replica numbers.
  • The examples demonstrated equivalence among methods in appropriate limits, with log-probability converging to the optimum as 1/N for replicas and 1/n for EROS.
  • Changing θ is equivalent in the optimal distributions to uniformly scaling all squared Gaussian errors, enabling efficient θ assessment through relative-entropy estimates.
  • The procedure assumes a reference distribution, while combining multiple potential energy functions may improve coverage of relevant phase space.The paper also notes that Gaussian error models may be replaced by more general likelihood functions.
  • The paper leaves the relation between Bayesian approaches and minimal-set orthogonal refinement incompletely understood, including their limiting behavior for large sample sizes.

Appendix A: Adaptive sampling of optimal Bayesian distribution without replicas

The appendix develops an adaptive, replica-free algorithm for sampling the optimal Bayesian ensemble distribution while incorporating measurement errors. The generalized forces are adjusted from running observable averages, and convergence is established analytically for a Gaussian example and observed numerically with Monte Carlo sampling.

  • Observable probability densities are defined by integrating out all other configuration-space degrees of freedom.
  • The adaptive method determines generalized forces self-consistently rather than requiring exact agreement with observed values, thereby incorporating measurement errors.The forces are adjusted on the fly according to running averages of the observables.
  • The trajectory evolves under a time-dependent potential that couples the configuration observables to the adaptive generalized forces.Extending the phase space to include configurations and forces yields a Markovian formulation with a Liouville-type evolution equation.
  • For the Gaussian example with the mean as observable and overdamped diffusion, adaptive sampling globally converges to the optimal Bayesian ensemble distribution.For the corresponding Monte Carlo example in Fig. 2, convergence was observed numerically.
  • The adaptive generalized forces can combine an initial guess with evolving observable averages through a time-dependent weight that decreases to zero.One example is an exponentially decaying weight, w(t) = exp(−t/t0).
Loading 1509.04447v2…