Source-linked AI summary

Bayesian model averaging: A systematic review and conceptual classification

Tiago M. Fragoso, Francisco Louzada Neto

arXiv:1509.08864v1stat.MEstat.AP

TL;DR

Model selection can ignore model uncertainty, producing overconfident inferences and riskier decisions. This review describes BMA’s model-averaging framework, classifies its literature, and finds limited methodological development alongside identifiable gaps in model-prior and evidence-estimation practice.

  • Problem

    Selecting one model can ignore model uncertainty, leading to overconfident inferences and riskier decision making.

  • Method

    The review develops a conceptual classification scheme and analyzes key components and research trends in the BMA literature.

  • Results

    Methodological developments remained limited beyond seminal late-1990s work, with applications concentrated mostly on (generalized) linear regression models.

  • Takeaways & Limitations

    BMA provides a flexible account of model uncertainty, while evidence estimation and model-prior specification remain important areas represented in the reviewed literature.

  • Takeaways & Limitations

    The search covered peer-reviewed articles in digitally available periodicals listed in four databases, so conference proceedings, dissertations, theses, books, and non-listed titles may have been overlooked.

Abstract

from arXiv · show

Bayesian Model Averaging (BMA) is an application of Bayesian inference to the problems of model selection, combined estimation and prediction that produces a straightforward model choice criteria and less risky predictions. However, the application of BMA is not always straightforward, leading to diverse assumptions and situational choices on its different aspects. Despite the widespread application of BMA in the literature, there were not many accounts of these differences and trends besides a few landmark revisions in the late 1990s and early 2000s, therefore not taking into account any advancements made in the last 15 years. In this work, we present an account of these developments through a careful content analysis of 587 articles in BMA published between 1996 and 2014. We also develop a conceptual classification scheme to better describe this vast literature, understand its trends and future directions and provide guidance for the researcher interested in both the application and development of the methodology. The results of the classification scheme and content review are then used to discuss the present and future of the BMA literature.

1. INTRODUCTION

BMA addresses model uncertainty by assigning posterior probabilities to multiple models, enabling model selection, combined estimation, and prediction. Its application remains technically difficult because model priors and model evidence often require nontrivial choices or approximations.

  • 1. INTRODUCTION: Selecting one adequate model can produce overconfident inferences and riskier decisions by ignoring uncertainty about alternative model assumptions.The paper motivates combining or selecting multiple models to account for this uncertainty.
  • 1. INTRODUCTION: BMA extends Bayesian inference by modeling uncertainty over both parameters and competing models.Posterior parameter and model probabilities are obtained using Bayes’ theorem.
  • 1. INTRODUCTION: Posterior model probabilities support direct model selection and weighting of posterior distributions for combined estimates or predictions.The weighted average uses each model’s posterior probability.
  • 1. INTRODUCTION: BMA predictions have lower risk under a logarithmic scoring rule than predictions from a single model.This result applies to the practice of combining model-based predictions.
  • 1. INTRODUCTION: Applying BMA requires specifying model priors and calculating model evidence, which is usually nontrivial outside simple conjugate settings.Evidence often lacks a closed form and must be approximated.
  • 1. INTRODUCTION: The review fills a gap after earlier landmark reviews by classifying BMA developments and summarizing research findings and trends.It also aims to guide researchers applying BMA to complex models.

2. SURVEY METHODOLOGY

The study combines systematic review criteria with content analysis to examine explicitly documented BMA research published from 1996 to 2014. Searches across four databases yielded 587 eligible articles for classification.

  • 2. SURVEY METHODOLOGY: The authors conducted a content analysis within a systematic review framework using objective criteria for identifying and reporting relevant BMA literature.The approach was intended to identify contributions, trends, practices, and research possibilities.
  • 2. SURVEY METHODOLOGY: Four databases—Scopus, ScienceDirect, Web of Science, and MathSciNet—were searched for BMA publications.The search targeted the phrase “Bayesian Model Averaging.”
  • 2. SURVEY METHODOLOGY: The search covered publications from 1996–2014 using title, abstract, keyword, topic, or broad database fields, depending on the source.The period was chosen to cover literature not included in previous reviews while retaining seminal works.
  • 2. SURVEY METHODOLOGY: Articles had to be English-language, peer-reviewed journal articles available online and explicitly employ BMA.Conference proceedings, theses, dissertations, books, and papers merely mentioning BMA were excluded.
  • 2. SURVEY METHODOLOGY: 587 articles remained for classification after duplicates and eligibility exclusions were removed.The process produced 703 articles for detailed investigation and excluded 116 under the second criterion.
  • 2. SURVEY METHODOLOGY: The final dataset was classified by publication year, authors, journal, and responses to eight CCS items.These categories supported systematic description of the selected literature.

3. A CONCEPTUAL SCHEME FOR BMA

The paper develops a Conceptual Classification Scheme tailored to BMA literature and refines its categories while reviewing the selected articles. The scheme is intended to organize methodological characteristics and make research developments easier to query.

  • 3. A CONCEPTUAL SCHEME FOR BMA: The authors used content analysis to identify applications, trends, research directions, and opportunities in the selected BMA literature.The analysis was conducted without preconceived notions through conventional content analysis.
  • 3. A CONCEPTUAL SCHEME FOR BMA: The CCS was adapted from an earlier classification scheme to represent characteristics relevant to Bayesian Model Averaging.Categories were refined as all articles were reviewed.
  • 3. A CONCEPTUAL SCHEME FOR BMA: The eight-item CCS records key methodological and formulation characteristics of each article.The possible responses are presented in the paper’s Table 1 and accompanying descriptions.
  • 3. A CONCEPTUAL SCHEME FOR BMA: An article’s CCS classification can support more efficient searches for developments and applications of BMA.The scheme provides researchers with key characteristics for navigating the literature.

3.1 Usage

The review classifies BMA usage into major application and discussion types because BMA supports both model choice and model-averaged estimation or prediction. It distinguishes these uses from theoretical extensions and review articles.

  • 3.1 Usage: BMA usage was organized into five main categories after classifying each eligible article by application.The classification reflects the broad range of problems addressed by BMA.
  • 3.1 Usage: Posterior model probabilities provide an interpretable model-choice criterion without bookkeeping over parameter counts or penalty types.The authors therefore identified articles using BMA for model choice.
  • 3.1 Usage: Posterior model probabilities can weight estimates across models, accounting for model uncertainty and potentially reducing overall risk.The quantity of interest may be a parameter common to all models or a future observation.
  • 3.1 Usage: The review distinguishes joint estimation from prediction according to whether the averaged quantity is a common parameter or future data point.Both uses are based on the same model-averaging expression but have distinct applications.
  • 3.1 Usage: The classification also includes conceptual discussions and review articles because BMA’s application involves substantial technical literature and field-specific retrospectives.These categories capture theoretical aspects and reviews of BMA or related applications.

3.2 Field of application

The review groups BMA applications into four broad fields, using a general classification to identify research trends rather than provide a detailed taxonomy. These fields span statistical and machine-learning work, biological and life sciences, economics and humanities, and engineering and physical sciences.

  • Classification rationale: The authors use a deliberately vague classification because their goal is to summarize research trends, not develop a detailed taxonomy.They also cite limited expertise for discriminating among applied subfields.
  • Application fields: The four application fields are Statistics and Machine Learning, Biological and Life Sciences, Economics and Humanities, and Engineering and Physical Sciences.The Statistics and Machine Learning category combines statistical modeling, theoretical developments, and machine-learning applications.

3.3 Model priors

Model-prior choices are a central challenge in BMA because assigning probabilities across models is not obvious, and authors often leave priors unspecified. The review identifies vague, elicited, literature-based, and unavailable model-prior categories.

  • Prior specification: Model-prior elicitation is difficult because a probability measure over the model space is not obvious in BMA.The authors suggest this difficulty may be reflected in many papers’ failure to state priors explicitly.
  • Prior categories: A vague prior treats all models as equally likely a priori and lets the observed data provide the information.This corresponds to π(M_l) ∝ 1 across the K models.
  • Prior categories: Elicited priors incorporate expert opinions or problem-specific characteristics when a uniform prior is not desirable.
  • Prior categories: The review classifies inherited convenient choices as literature priors and unstated or unsupported model priors as not available.

3.4 Evidence estimation

Estimating model evidence is difficult because it often requires integrating complicated, high-dimensional functions without a closed form. The review organizes proposed solutions into five categories, including Monte Carlo, MCMC, ratio-of-densities, analytical approximations, and closed-form cases.

  • Evidence-estimation challenge: Model evidence estimation is non-trivial because it can involve high-dimensional integrands with complex support sets that make integration infeasible.
  • Stochastic approximations: Monte Carlo and importance-sampling methods approximate evidence using weighted averages of random parameter samples and converge by the Strong Law of Large Numbers.Bridge sampling is treated within the same Monte Carlo category.
  • MCMC methods: Importance Sampling estimators converge as sample size increases, but the importance function must be tuned to avoid unbounded variance.The harmonic-mean estimator is described as a special case using the prior as the importance density.
  • MCMC methods: MCMC-based methods estimate model posteriors from sampled model visits, including Gibbs or Metropolis sampling, reversible-jump methods, and stochastic searches.The review classifies MCMC-based methods into a single category.
  • Analytical approximations: Ratio-of-densities methods estimate evidence at a selected parameter point, while Laplace approximations and BIC use analytical or asymptotic approximations.The Laplace approximation is O(n^-1) under regularity conditions, whereas BIC has approximation error O(1).
  • Closed-form cases: Some articles compute evidence in closed form under particular model or prior assumptions and are classified as non-applicable for general evidence approximation.

3.5 Dimensionality

BMA faces dimensionality problems because regression with p covariates yields 2^p possible subsets, making exhaustive model evaluation impractical even for moderate p. The literature responds with model-space reduction and stochastic MCMC searches.

  • Problem: Regression variable selection has 2^p possible models when interaction terms are excluded, causing model spaces to grow geometrically with p.This growth can preclude exhaustive investigation of all models.
  • Dimensionality reduction: Dimensionality reduction includes Leaps and Bounds, which obtains parsimonious regression models without exhaustive search by exploiting linear-model relationships.Its selection criterion is residual sum of squares.
  • Dimensionality reduction: Occam’s Window retains models whose posterior probabilities remain close to that of the most likely model.The cutoff is commonly set to 20, corresponding to a 0.05-style filtering threshold.
  • Stochastic searches: MCMC searches address dimensionality by exploring the full model space and estimating posterior model probabilities from model-visit frequencies.
  • Stochastic searches: SSVS updates indicator variables whose configurations represent distinct models, while RJMCMC and MC3 move between models using Markov-chain procedures.The review groups these MCMC-based searches into one category.
  • Non-applicable cases: When dimensionality is not an issue, some applications fit all models and perform BMA posteriorly without proposing a mitigation method.These cases are classified as not applicable.

3.6 Markov Chain Monte Carlo methods

MCMC methods enabled inference for complex likelihood and prior structures and became a default approach through accessible software, including BUGS and JAGS.

  • MCMC enabled inference with complex likelihood and prior structures using straightforward posterior outputs grounded in solid theory.
  • MCMC methods became the default approach for many applied Bayesian problems after software dissemination.
  • BUGS and JAGS, commonly integrated with R, broadened access to robust MCMC implementations and Bayesian inference.

3.7 Simulation studies

The review classified simulation studies as uses of generated model simulations or artificial datasets, covering investigations of averaging properties, predictive power, physical systems, and diverse-model predictions.

  • Simulation studies were identified through model-generated simulations or artificial datasets.
  • These studies investigated averaging characteristics and properties such as predictive power in best-case scenarios.
  • Simulations also served to emulate physical systems and generate predictions from diverse models for averaging.

3.8 Data-driven validation

The review distinguished simple data splitting, K-fold or leave-one-out cross-validation, and Bayesian posterior predictive checks as data-driven validation approaches after BMA.

  • Simple cross-validation splits data into disjoint fitting and validation sets.
  • K-fold cross-validation fits models on K−1 subsets and validates them on each held-out subset in turn.
  • Leave-one-out cross-validation was grouped with K-fold cross-validation as a sophisticated validation procedure.
  • Posterior Predictive Checks compare an observed-data test statistic with replicated data generated from the posterior distribution.

4. RESULTS AND DISCUSSION

The review of 587 BMA articles shows rapid post-2005 growth, broad disciplinary diffusion, and predominant use for model choice, especially regression variable selection. Prediction and estimation were also substantial, while methodological work and reviews were smaller categories.

  • Descriptive statistics: After 2005, BMA publications increased substantially, alongside accessible computational tools, computer power, and ready-to-use MCMC software.
  • Usage: 231 works, almost 40% of the dataset, used BMA for model choice, predominantly variable selection in regression models.
  • Usage: Model-choice applications included Occam’s Window, Leaps and Bounds, explicit conjugate-prior searches, MC3, Reversible Jump MCMC, and SSVS.
  • Usage: BMA was used for combined prediction through fitting models, calculating evidences, and combining model-specific predictive distributions.
  • Usage: 111 articles, around 19%, used BMA for combined estimation of quantities modeled across multiple candidate models.
  • Methodological and review work: 65 articles, around 11%, were conceptual or methodological contributions, while 18 review articles accounted for less than 3% of the dataset.
  • Fields of application: 201 articles came from Life Sciences and Medicine, 149 from Humanities and Economy, 146 from Physical Sciences and Engineering, and 92 from Statistics and Machine Learning.

5. CONCLUDING REMARKS

The review classifies 587 BMA articles from 1996–2014 and uses the results to identify limitations and future directions in the literature. It also notes important scope limitations in the review itself, including restricted databases, publication types, and search terms.

  • 587 articles were classified using a conceptual scheme covering BMA literature from 1996–2014.The review spans diverse publications and applications.
  • The review may omit developments from conference proceedings, dissertations, theses, books, non-indexed periodicals, and articles missed by its database queries.
  • The authors present the review as a source of insights into current BMA research and indications of future developments, without claiming exhaustiveness.
  • The review identifies limited methodological development beyond seminal late-1990s work, with applications concentrated mainly in regression and Bayesian-network settings.Evidence estimation often relies on convenient likelihoods, conjugate priors, or BIC approximations, while newer MCMC developments were largely absent.
  • Validation is important because BMA combines models with potentially distinct assumptions and conclusions, and unvalidated use may recreate overconfidence.
Loading 1509.08864v1…