Source-linked AI summary

Approximate Bayesian Computational methods

Jean-Michel Marin, Pierre Pudlo, Christian P. Robert, Robin Ryder

arXiv:1101.0955v2stat.CO

TL;DR

ABC addresses Bayesian inference when likelihoods are unavailable or computationally impractical. This survey synthesizes ABC foundations and methodological extensions, including calibration, sequential improvements, post-processing, and model choice. It concludes that ABC model choice generally lacks convergence guarantees and theoretical support, motivating further empirical assessment.

  • Problem

    ABC is needed for Bayesian models whose likelihoods are unavailable or impractical to evaluate, while summary-statistic choice remains a fundamental difficulty.

  • Method

    The survey reviews ABC’s foundations, calibration, sequential improvements, post-processing, and model-choice methods.

  • Results

    The survey reports that ABC model-choice performance can remain inaccurate, with an MA(2) posterior probability estimate of 0.72 versus a true value of 0.95 at the 0.01% quantile.

  • Takeaways & Limitations

    ABC model choice has exact ε = 0 simulation for Gibbs random fields through cross-model sufficient statistics, but this property is specific to exponential families.

  • Takeaways & Limitations

    General ABC model choice lacks convergence guarantees and theoretical support for Bayes factors and posterior model probabilities.

Abstract

from arXiv · show

Also known as likelihood-free methods, approximate Bayesian computational (ABC) methods have appeared in the past ten years as the most satisfactory approach to untractable likelihood problems, first in genetics then in a broader spectrum of applications. However, these methods suffer to some degree from calibration difficulties that make them rather volatile in their implementation and thus render them suspicious to the users of more traditional Monte Carlo methods. In this survey, we study the various improvements and extensions made to the original ABC algorithm over the recent years.

1 Introduction

ABC addresses Bayesian inference when likelihoods are unavailable, incomplete, or impractical to evaluate by replacing exact computation with simulation-based approximations. The survey introduces ABC and outlines its scope and assessment through methodological developments and time-series examples.

  • Likelihoods can be unavailable because they lack a closed-form expression or are too expensive to calculate.
  • Latent-variable integration can make generic MCMC methods impractical because data augmentation increases dimension and may yield poor convergence.
  • Unknown normalizing constants create another class of likelihood problems, including Gibbs random fields for spatially correlated data.
  • ABC provides an almost automated likelihood-free approach for models that are intractable but simulable, extending beyond earlier approximation methods.
  • The survey covers ABC foundations, calibration, sequential improvements, post-processing, and model choice, assessing approximation effects with MA(1) and MA(2) posteriors.

2 Genesis of the ABC approach and justifications

ABC begins as a rejection sampler that simulates parameters and data, accepting parameters whose simulated summaries are sufficiently close to the observed summaries. The section develops this construction, illustrates its approximation, and presents MCMC, noisy, and filtering extensions.

  • Genesis of the ABC approach: The original ABC algorithm draws θ from the prior, simulates data under θ, and accepts θ when simulated and observed samples are nearly identical.
  • Genesis of the ABC approach: For continuous data, ABC replaces exact matching with a distance between summary statistics and a tolerance ε.
  • Genesis of the ABC approach: The basic approximation depends on choosing a representative summary statistic and a sufficiently small tolerance.
  • Genesis of the ABC approach: In the MA(2) illustration, autocovariance summaries produce an ABC sample that departs from the exact posterior even at a 0.1% acceptance threshold.
  • Extensions: MCMC-ABC targets the approximate posterior without likelihood evaluation, but its calibration depends on the summary statistic, tolerance, and distance.
  • Extensions: Noisy ABC replaces loose acceptance with exact inference under a kernel-convolved target, while ABC filtering allows complex models and avoids accumulating approximation errors along observations.

3 Calibration of ABC

ABC calibration depends critically on summary statistics, distances, and tolerance levels. The survey reviews sequential selection and error-assessment strategies, illustrating that smaller tolerances improve approximation but do not eliminate discrepancies from the true posterior.

  • Summary statistics and calibration: ABC studies have examined sequential inclusion of summary statistics, but some methods still depend on an approximation factor calibrated before execution.
  • Summary statistics and calibration: The choice of distance, summary statistics, and calibration factors is paramount to the success of ABC approximation.
  • Calibration perspectives: The survey contrasts calibration perspectives that target posterior approximation with approaches designed for inferential prediction and well-calibrated quantities.
  • Tolerance threshold: Ratmann et al. treat tolerance as an additional model parameter and use its marginal posterior to assess model fit.
  • Tolerance threshold: The ABCµ approach has a delicate Bayesian interpretation because the prior on ε may significantly influence model assessment.
  • Tolerance threshold: Smaller tolerance levels improve ABC approximation in the MA(2) example, although the true marginal densities are never reached, particularly for θ2.

4 Sequential improvements

Sequential ABC methods improve efficiency by adapting particles, kernels, importance weights, or tolerance thresholds across iterations. The survey distinguishes methods with genuine importance-sampling foundations from approaches that can introduce bias or retain computational limitations.

  • Backward kernels and SMC: Replacing the likelihood with an ABC indicator introduces bias in the approximation to the posterior in the method discussed by Del Moral et al.
  • Importance sampling: ABC-PMC constructs target approximations from earlier simulations, automatically scales kernels, and uses decreasing tolerance thresholds.
  • Sequential Monte Carlo: ABC-SMC generates particle populations through sequential samples, Markov transition kernels, and importance weights while progressively reducing tolerance.
  • Empirical comparisons: In a single-dataset comparison with fixed random-walk variances, ABC-SMC outperformed ABC-MCMC, while the simulation’s prior might have been overly concentrated.
  • Backward kernels and SMC: Backward kernels can simplify importance weights and remove their dependence on the unavailable likelihood within an ABC-SMC framework.
  • Adaptive methods: Adaptive particle methods use previous populations to choose proposal kernels and tolerance levels, but repeated MCMC steps can reduce the appeal of O(N) methods.

5 Post-processing of ABC output

ABC post-processing improves inference by correcting simulated parameters after sampling, allowing larger tolerances while reducing distortion from tolerance-based truncation. The section also describes nonlinear and inverse-regression extensions to this correction framework.

  • Local linear regression: Local linear regression treats ABC post-processing as conditional density estimation, shrinking simulated parameters toward the observed summary statistic to permit larger ε.The simulation process remains unchanged; only the ABC output is analyzed differently.
  • Local linear regression: In the MA(2) example, the Beaumont correction brings θ1 closer to the true posterior at the 0.1% tolerance, while θ2 estimates are identical.The comparison uses the first two autocovariances as summaries and contrasts regular ABC with the post-processed estimate.
  • Local linear regression: At the 20% tolerance, local regression produces results close to those obtained with the 0.1% tolerance, strongly attenuating ε-induced truncation.The figure compares regular ABC and Beaumont-corrected distributions at the larger tolerance.
  • Nonlinear regression: Blum and François generalize Beaumont’s post-processing by replacing local linear regression with heteroskedastic nonlinear regression estimated by a one-hidden-layer neural network.The nonlinear approach estimates both the conditional mean and variance.
  • Inverse regression: Inverse-regression approaches model summary statistics given parameters rather than parameters given summaries, using a Gaussian linear approximation followed by a Laplace-type approximation.Unlike Beaumont et al., accepted parameters are weighted similarly rather than clearly shrunk according to summary-statistic distance.

6 ABC and model choice

ABC model choice estimates posterior model probabilities through acceptance frequencies, but their accuracy depends critically on summary-statistic sufficiency. Exponential-family settings can provide exact model-choice simulation, whereas the MA example and broader discussion show substantial unresolved limitations.

  • ABC model-choice procedure: ABC-MC samples a model index and parameters from their priors, simulates data, and accepts proposals whose summary-statistic distance is below ε.
  • ABC model-choice procedure: Posterior model probabilities are estimated by the acceptance frequency for each model, with logistic regression providing a more stable alternative.
  • Empirical model choice: For MA(2) data, reducing ε to the 0.01% quantile raises the estimated MA(2) posterior probability only to 0.72, versus the true value 0.95.
  • Gibbs random fields: For Gibbs random fields, a cross-model sufficient-statistic vector exists because of their exponential-family structure, enabling exact ABC model-probability simulation at ε = 0.
  • Gibbs random fields: Concatenating model-specific sufficient statistics is generally not sufficient for model choice outside exponential families, especially when the full data are too demanding.
  • General issues: ABC model-choice Bayes factors currently lack convergence guarantees, so the authors advise empirical model-fit assessments rather than treating estimated probabilities as exact.

7 Discussion

ABC broadens Bayesian inference to otherwise unavailable models, while calibration, theoretical guarantees, summary-statistic selection, and computational scalability remain open issues. Sequential methods and regression post-processing improve efficiency, but model choice remains limited.

  • ABC enables inference for a wide class of models that would otherwise be unavailable, and recent calibration advances can make its approximation useful in many situations.
  • Sequential techniques and post-processing regression can greatly improve the efficiency of ABC outputs.
  • ABC model choice is limited because its approximation cannot currently be guaranteed valid, even when a large collection of summary statistics is used.
  • Current convergence results require the tolerance to approach zero or the sample size to approach infinity, leaving finite-tolerance error bounds unresolved.
  • The construction and selection of summary statistics remains highly empirical, motivating automated approaches based on data analysis and approximate sufficiency.
  • Large datasets and complex models can make pseudo-data simulation impossible, motivating methods that simulate summary statistics directly.
Loading 1101.0955v2…