Source-linked AI summary

Bayesian Post-Processor and other Enhancements of Subset Simulation for Estimating Failure Probabilities in High Dimensions

Konstantin M. Zuev, James L. Beck, Siu-Kui Au, Lambros S. Katafygiotis

arXiv:1110.3390v1stat.COstat.AP

TL;DR

The paper tackles estimation of small failure probabilities in high-dimensional reliability problems, where direct integration is difficult. It enhances Subset Simulation through Modified Metropolis tuning, analysis of p0, and the Bayesian SS+ post-processor. The paper recommends acceptance rates of 30%–50% for MMA and finds that p0 values from 0.1 to 0.3 have similar efficiency.

  • Problem

    Estimating failure probabilities over high-dimensional uncertain parameter spaces is difficult because the required reliability integral is hard to evaluate numerically.

  • Method

    The paper enhances Subset Simulation by tuning MMA, analyzing the conditional failure probability p0, and developing the Bayesian SS+ post-processor.

  • Results

    Choosing p0 between 0.1 and 0.3 leads to similar efficiency without fine tuning when Subset Simulation is properly implemented.

  • Takeaways & Limitations

    SS+ represents the unknown failure probability with a posterior PDF that quantifies uncertainty and can be used in risk analyses or for point estimation.

  • Takeaways & Limitations

    The SS+ posterior approximation treats dependent MCMC samples as if they were independent, with some reduction in efficiency from sample correlation.

Abstract

from arXiv · show

Estimation of small failure probabilities is one of the most important and challenging computational problems in reliability engineering. The failure probability is usually given by an integral over a high-dimensional uncertain parameter space that is difficult to evaluate numerically. This paper focuses on enhancements to Subset Simulation (SS), proposed by Au and Beck, which provides an efficient algorithm based on MCMC (Markov chain Monte Carlo) simulation for computing small failure probabilities for general high-dimensional reliability problems. First, we analyze the Modified Metropolis algorithm (MMA), an MCMC technique, which is used in SS for sampling from high-dimensional conditional distributions. We present some observations on the optimal scaling of MMA, and develop an optimal scaling strategy for this algorithm when it is employed within SS. Next, we provide a theoretical basis for the optimal value of the conditional failure probability $p_0$, an important parameter one has to choose when using SS. Finally, a Bayesian post-processor SS+ for the original SS method is developed where the uncertain failure probability that one is estimating is modeled as a stochastic variable whose possible values belong to the unit interval. Simulated samples from SS are viewed as informative data relevant to the system's reliability. Instead of a single real number as an estimate, SS+ produces the posterior PDF of the failure probability, which takes into account both prior information and the information in the sampled data. This PDF quantifies the uncertainty in the value of the failure probability and it may be further used in risk analyses to incorporate this uncertainty. The relationship between the original SS and SS+ is also discussed

1 Introduction

The paper addresses the difficulty of estimating failure probabilities over high-dimensional uncertain spaces by enhancing Subset Simulation. Its enhancements target Modified Metropolis sampling, the conditional failure probability p0, and Bayesian uncertainty quantification.

  • Motivation: Failure-probability estimation is challenging because reliability problems involve integrals over high-dimensional uncertain parameter spaces.The dimension can be large in dynamic reliability problems because stochastic input histories are discretized in time.
  • Motivation: Subset Simulation is an efficient algorithm for computing small failure probabilities in general high-dimensional reliability problems.Prior theory and numerical examples found it more computationally efficient than standard Monte Carlo Simulation for small failure probabilities.
  • Enhancements: The paper develops an optimal scaling strategy for the Modified Metropolis algorithm used within Subset Simulation.The strategy is based on observations about MMA scaling and aims to improve Markov-chain convergence to stationarity.
  • Enhancements: The paper presents a method for choosing the intermediate failure probabilities that govern Subset Simulation efficiency.These probabilities are equivalent to the sequence of intermediate threshold values.
  • Bayesian extension: SS+ models the unknown failure probability as a stochastic variable and produces its posterior PDF using prior information and samples generated by Subset Simulation.The posterior PDF quantifies the relative plausibility of possible failure-probability values and can support risk analyses or point estimation.

2 Subset Simulation

Subset Simulation replaces direct estimation of a very small failure probability with sequential estimation of larger conditional probabilities. It uses Modified Metropolis Markov chains to sample the required conditional distributions, while proposal spread and intermediate failure domains determine efficiency.

  • Monte Carlo motivation: 26?
  • Subset Simulation formulation: Subset Simulation represents a small failure probability as a product of larger conditional probabilities over nested intermediate failure domains.The original problem is thereby replaced by a sequence of intermediate probability-estimation problems.
  • Subset Simulation formulation: Conditional probabilities at intermediate levels can be estimated from samples drawn from the corresponding conditional distributions.The first level can use standard Monte Carlo sampling, whereas higher levels require sampling from π(·|Fj−1).
  • Modified Metropolis algorithm: Modified Metropolis uses coordinate-wise univariate proposals to generate candidate states and targets π(·|F) as the stationary distribution of its Markov chain.The algorithm accepts or rejects candidates after checking whether they lie in the failure domain.
  • Modified Metropolis algorithm: MCMC samples are correlated, so proposal distributions affect estimator efficiency through the chain’s dependence and convergence.Both overly small and overly large proposal spreads can increase dependence or reduce acceptance rates.
  • Subset Simulation algorithm: Intermediate failure domains are defined from a performance function g and a prescribed critical threshold for unacceptable system performance.These domains provide the nested sequence required by Subset Simulation.

3 Tuning of the Modified Metropolis algorithm

This section examines how to tune the Modified Metropolis algorithm within Subset Simulation, focusing on proposal scaling, acceptance rates, and practical heuristics. Optimal scaling can improve estimation efficiency, but its problem-specific parameters are difficult to determine.

  • Proposal scaling: MMA’s proposal distributions control how quickly its Markov chains explore the parameter space and converge to stationarity within Subset Simulation.The efficiency and accuracy of SS directly depend on the ergodic properties of these chains.
  • Proposal scaling: Optimal scaling tunes algorithm parameters to maximize the convergence speed of the resulting Markov chain.For the original Metropolis algorithm, high-dimensional Gaussian targets have an approximately 23% optimal acceptance rate.
  • Proposal scaling: For i.i.d. target components, the optimal Gaussian proposal variance depends on dimension and the one-dimensional density’s smoothness.For a standard Gaussian in one dimension, the cited optimal proposal spread is approximately σ ≈ 2.4.
  • Practical tuning: A practical heuristic is to tune the proposal variance so the average acceptance rate is roughly 25%.The paper notes that this heuristic is believed to remain robust under various perturbations of the target distribution.
  • Results: When pF = 10^-6 and N = 1000, optimal scaling reduces the coefficient of variation to approximately 80% of the reference value in both examples.The comparison is between σj = σopt_j and the reference choice σj = 1.
  • Practical tuning: Although γj is relatively flat around the optimal acceptance rate, finding problem-specific optimal spreads can be computationally expensive.The authors therefore seek an easy-to-implement heuristic that provides near-optimal behavior.

4 Optimal choice of conditional failure probability p0

This section derives a basis for choosing the conditional failure probability p0 in Subset Simulation. The analysis identifies an approximately optimal value while showing that performance is relatively insensitive across a broader practical range.

  • Role of p0: The choice of p0 controls how many intermediate failure domains are used and the number of samples required at each level.Smaller p0 means fewer intermediate levels but requires more samples to estimate the smaller conditional probabilities accurately.
  • Optimization: The proposed criterion chooses p0 to minimize the estimator’s coefficient of variation for a fixed total number of samples.The analysis assumes equal conditional-probability estimates and the same sample count at each level.
  • Optimization: The theoretical analysis gives an approximately optimal conditional failure probability of p0 ≈ 0.2.For fixed target failure probability and total sample count, the relevant factor determining the coefficient of variation is minimized near this value.
  • Practical range: Choosing 0.1 ≤ p0 ≤ 0.3 produces practically similar efficiency because the coefficient of variation is relatively insensitive around its optimum.The reported trend shape is invariant with respect to pF, NT, and the average correlation factor because their effects are multiplicative.

5 Bayesian Post-Processor for Subset Simulation

The Bayesian post-processor SS+ treats the conditional probabilities underlying Subset Simulation as uncertain quantities and uses sampled data to produce a posterior distribution for the failure probability. The original SS estimate remains connected to SS+ as its MAP estimate, while the posterior quantifies uncertainty for risk-related analyses.

  • Bayesian formulation: The post-processor specifies priors, updates them with level-specific SS data, and combines the resulting posterior distributions through p_F = product of p_j.The posterior for p_F is obtained from the posterior distributions of the conditional probabilities.
  • Bayesian formulation: SS+ models each conditional probability p_j and the overall failure probability p_F as stochastic variables, rather than treating their estimated values as fixed.The failure probability is constant but unknown because its defining integral cannot be evaluated exactly.
  • Posterior construction: For the first conditional level, Bernoulli failure indicators yield a beta posterior Be(n_1 + 1, N − n_1 + 1) under the uniform prior.At higher levels, analogous beta expressions are used as approximations because MCMC samples are dependent.
  • Posterior construction: The higher-level approximation is supported by multiple Markov chains with different initial seeds, which makes it more accurate than using a single chain.The approximation nevertheless treats the samples as independent and therefore neglects their MCMC dependence.
  • Relationship between SS and SS+: The original SS estimate is the MAP estimate of the SS+ posterior and accurately approximates the mean of its beta approximation.The c.o.v. of the exact and approximate SS+ distributions coincide and can be computed from their first two moments.
  • Practical use: The full posterior PDF can replace a point estimate in life-cycle cost analysis and decision making under risk to incorporate uncertainty in p_F.An expected loss can be calculated when a performance loss function depends on the failure probability.

6 Illustrative Examples

The paper demonstrates SS+ on linear and nonlinear reliability problems, showing that it produces posterior failure-probability distributions whose uncertainty can be quantified from SS samples. The examples also show narrower posterior PDFs with more samples.

  • SS+ is demonstrated on a linear reliability problem and an elasto-plastic structure subjected to strong seismic ground motion.
  • Linear reliability problem: For the linear problem, SS estimates pF = 1.057 × 10^-3, close to the exact value pF = 10^-3.The example uses dimension d = 10^3 and three conditional levels with N = 10^3 samples at each level.
  • Linear reliability problem: The linear example gives an SS+ posterior mean of 1.064 × 10^-3 and c.o.v. of 0.16 around the SS estimate.The posterior is described as narrowly focused around the SS estimate.
  • Elasto-plastic structure subjected to ground motion: For the nonlinear structure, Monte Carlo simulation estimates pF = 8.9 × 10^-3, while SS+ is evaluated with N = 500, 1000, and 2000 samples per level.The structure is modeled as a two-dimensional six-story steel frame under uncertain seismic excitation, with failure defined by an interstory drift threshold.
  • Elasto-plastic structure subjected to ground motion: SS+ posterior c.o.v. values are 0.190, 0.134, and 0.095 for N = 500, 1000, and 2000, respectively.The posterior becomes more narrowly focused around the SS estimate as the number of samples increases.
  • Elasto-plastic structure subjected to ground motion: In SS+, the posterior coefficient of variation measures uncertainty based on the generated samples.

7 Conclusions

The conclusions present three enhancements to Subset Simulation: MMA tuning, a theoretically supported range for p0, and the Bayesian SS+ post-processor. SS+ replaces a single estimate with a posterior PDF that combines prior and sampled information.

  • The paper enhances Subset Simulation, an algorithm for computing failure probabilities in general high-dimensional reliability problems.
  • MMA scaling: For MMA within SS, choose σ_j so the acceptance rate ρ_j lies between 30% and 50% at simulation level j ≥ 1.This is presented as a nearly optimal scaling strategy.
  • Conditional failure probability: Any conditional failure probability p0 ∈ [0.1, 0.3] gives similar efficiency when SS is implemented properly, so fine tuning is unnecessary.
  • Bayesian post-processing: SS+ models the unknown failure probability pF as a stochastic variable and produces its posterior PDF instead of a single real-number estimate.
  • Bayesian post-processing: The SS+ posterior incorporates prior information and SS samples, quantifies uncertainty in pF, and can be used in risk analyses; the original SS estimate is its most probable value.

Appendix

The appendix proves that the conditional target distribution π(·|F) is stationary for the Modified Metropolis chain. The proof uses reversibility, while uniqueness of the limiting distribution additionally requires aperiodicity and irreducibility.

  • The appendix states that π(·|F) is the stationary distribution of the Markov chain generated by MMA.
  • The proof represents the MMA transition kernel using transitions between distinct states, a stay probability, and component transition kernels.
  • Reversibility, or detailed balance, is used as a sufficient condition for π(·|F) to be stationary.
  • Symmetry of S_j and the identity b min{1, a/b} = a min{1, b/a} establish the required balance relation.
  • A stationary distribution is unique and limiting when the chain is aperiodic and irreducible; MMA aperiodicity follows from its nonzero probability of repeated samples.The appendix defines irreducibility as positive probability of entering any set assigned positive probability by the stationary distribution.
Loading 1110.3390v1…