Source-linked AI summary
Philosophy and the practice of Bayesian statistics
Andrew Gelman, Cosma Rohilla Shalizi
TL;DR
The paper challenges the view that Bayesian inference is inductive reasoning or rationality as such, given that successful Bayesian practice relies on model checking and revision. It examines priors, Bayesian consistency, posterior predictive checking, and applied examples, concluding that this practice accords better with sophisticated hypothetico-deductivism.
Problem
The paper addresses the disjunction between Bayesian inference understood as inductive reasoning and successful Bayesian practices of model building and model checking.
Method
The paper analyzes the roles of prior distributions, Bayesian consistency, posterior predictive simulation, model checking, and model revision, drawing on applied social-science examples.
Results
Successful Bayesian analyses rely on model checking and revision rather than calculating posterior probabilities that particular models are true.
Takeaways & Limitations
Bayesian statistical practice is better understood as incorporating falsification and model criticism than as exemplifying Bayesian confirmation theory alone.
Takeaways & Limitations
Posterior distributions may fail to provide reliable uncertainty representations in infinite-dimensional Bayesian nonparametric problems.
Abstract
from arXiv · showhide
A substantial school in the philosophy of science identifies Bayesian inference with inductive inference and even rationality as such, and seems to be strengthened by the rise and practical success of Bayesian statistics. We argue that the most successful forms of Bayesian statistics do not actually support that particular philosophy but rather accord much better with sophisticated forms of hypothetico-deductivism. We examine the actual role played by prior distributions in Bayesian models, and the crucial aspects of model checking and model revision, which fall outside the scope of Bayesian confirmation theory. We draw on the literature on the consistency of Bayesian updating and also on our experience of applied work in social science. Clarity about these matters should benefit not just philosophy of science, but also statistical practice. At best, the inductivist view has encouraged researchers to fit and compare models without checking them; at worst, theorists have actively discouraged practitioners from performing model checking because it does not fit into their framework.
1 The usual story—which we don’t like
The conventional philosophy presents Bayesian inference as inductive learning through changing posterior probabilities. The paper argues instead that successful Bayesian practice aligns better with hypothetico-deductive reasoning and model checking.
- The usual story—which we don’t like: The usual Bayesian story treats inference as learning general hypotheses from accumulating particular observations.Posterior probabilities summarize the changing degree of belief in competing models as evidence accumulates.
- The usual story—which we don’t like: The paper argues that Bayesian methods are no more inductive than other statistical approaches and are better understood hypothetico-deductively.It connects model checking with Mayo’s error-probing perspective despite that framework’s frequentist orientation.
- The usual story—which we don’t like: The paper examines this disjunction through philosophical discussion, theoretical results on Bayesian updating, and empirical social-science data analysis.The authors emphasize that social-science models are typically false, yet model fitting remains valuable.
- The usual story—which we don’t like: A central puzzle is that successful Bayesian modeling emphasizes model checking and continuous model expansion, unlike the received inductivist interpretation.The authors advocate treating the usual philosophical story as faulty rather than forcing applied practice into it.
2 The data-analysis cycle
Bayesian data analysis does not end with calculating a posterior. Analysts compare model implications with data, identify discrepancies, and revise the model through an iterative fitting and checking process.
- 2 The data-analysis cycle: A Bayesian model specifies a joint stochastic distribution for observed data, latent variables, and parameters, balancing representation against tractability.The likelihood is the joint distribution of the data as a function of parameters; Bayesian models factor the joint distribution into prior and complete-data likelihood.
- 2 The data-analysis cycle: Posterior calculation is only one stage: analysts simulate or otherwise compare fitted-model implications with observed and unused data to assess fit.Relevant checks include replicated datasets, residual patterns, and consistency with other observations.
- 2 The data-analysis cycle: Model discrepancies motivate expansions and changes rather than merely lowering confidence in a fixed set of competing models.Applied work involved fitting increasingly complex models while modifying priors and likelihoods in response to data-model comparisons.
- 2.1 Example: Estimating voting patterns in subsets of the population: In the voting example, a varying-intercept model failed to fit the data, motivating a varying-slope model and revealing geographic variation in rich-poor voting gaps.The model expansion enabled estimation across states, including states with relatively small samples through hierarchical partial pooling.
- 2.1 Example: Estimating voting patterns in subsets of the population: Further checks showed that even the more complex logistic model failed for ethnic-group analyses, leading to nonlinear and non-monotonic extensions.Bayesian inference supplied the computational framework for fitting richer models, while comparison of data and fit drove revision.
3 The Bayesian principal-agent problem
Bayesian updating is constrained by the model and prior support, especially when the model is false. The authors argue that checking Bayesian models is necessary to determine whether misspecification threatens scientific inferences.
- 3 The Bayesian principal-agent problem: Bayesian model selection does not solve the problem of learning from data when all available models are misspecified.Posterior probabilities of candidate models remain statements within the existing model space.
- 3 The Bayesian principal-agent problem: The Bayesian agent represents a methodological system with a prior, likelihood, and conditioning, while the actual scientist must manage imperfect and incomplete models.The philosophical formulation assumes an exact truth within a well-defined hypothesis space, unlike the authors’ experience of social-science modeling.
- 3 The Bayesian principal-agent problem: The truth must lie within prior support for standard posterior-concentration results, leaving the Bayesian agent unable to learn hypotheses excluded by its prior.Classical consistency results require both prior support for the truth and sufficiently rich information.
- 3 The Bayesian principal-agent problem: Under misspecification, the posterior concentrates on prior-supported parameter regions with the lowest divergence and highest expected likelihood.These regions can be closest approximations within the model without yielding accurate scientific parameters or a single limiting value.
- 3 The Bayesian principal-agent problem: Because misspecification can affect scientific inference from negligibly to profoundly, non-Bayesian model checking is used to assess whether the model is adequate for its purpose.The authors describe this checking as solving the Bayesian principal-agent problem.
4 Model checking
Bayesian model checking treats priors and likelihoods as assumptions whose predictive implications can be tested, revealing how models fail and guiding their revision. This practice supports a hypothetico-deductive understanding of Bayesian analysis, while posterior model probabilities and naive posterior uncertainty representations have important limitations.
- Testing model implications: Posterior predictive checks simulate replicated data from the fitted prior-and-likelihood model and compare those simulations with the observed data.These checks assess whether the observed dataset resembles typical realizations generated by the fitted model.
- Testing model implications: Extreme posterior predictive p-values indicate that the data violate, or nearly violate, probabilistic regularities implied by the model.Their logic generalizes classical p-values by averaging over the posterior distribution rather than using point estimates alone.
- The role of priors: The prior is part of the model, need not represent personal belief, and has testable implications for replicated or future data.This view treats the prior and likelihood as a compromise among scientific knowledge, mathematical convenience, and computational tractability.
- The role of priors: Bayesian modeling gains flexibility by stating assumptions clearly and checking their implications rather than requiring priors to match subjective beliefs or encompass all possible truths.The same checking principle applies to prior distributions and likelihood components.
- Model revision: Model checking aims to identify how a model fails, so detected mis-specifications can guide improvements while successful severe tests make related inferences more credible.Tests should target errors that would compromise particular inferences rather than merely establish that every model is false.
- Model comparison: Posterior model probabilities and model-comparison methods can be useful data-analytic tools, but the authors reject treating them as direct vehicles for scientific learning or unqualified degrees of belief.The paper emphasizes plots and predictive checks over Bayes factors or posterior probabilities of candidate models for falsification.
5 The question of induction
The paper rejects the idea that Bayesian inference is uniquely inductive, arguing that statistical inferences are deductively guaranteed only relative to probabilistic assumptions and substantive modeling conditions.
- The question of induction: Bayesian and classical statistical reasoning are better understood as hypothetico-deductive rather than as fundamentally different forms of inference.The authors contrast the received inductivist view of Bayesian statistics with an approach centered on hypotheses, implications, and testing.
- The question of induction: Inference from observations to general claims qualifies as induction only relative to assumptions linking observed data to broader conclusions.Those assumptions include probabilistic premises and, in applications such as sampling, material facts about the system under study.
- The question of induction: Bayesian updating is not sufficient or necessary for consistency, and consistency theorems provide deductively valid guarantees conditional on probabilistic assumptions.The paper notes that known proofs of Bayesian consistency either posit an equivalent consistent non-Bayesian procedure or rely on assumptions implying one.
- The question of induction: Conditioning follows deductively from accepted probability and decision-theoretic axioms, so axiomatization alone does not make Bayesian updating inductive.The authors also argue that many algorithmic learning procedures can be axiomatized.
- The question of induction: Statistical models support inductive-looking conclusions through deductions from sampling properties and other contingent facts, not through an independent axiom of induction.Random sampling can support learning about unsampled people when the relevant sampling properties and practical design conditions hold.
- The question of induction: Distribution-free learning results remain conditional on assumptions such as independent and identically distributed or stationary, mixing data-generating processes.Thus, “distribution free” does not mean free of all substantive premises.
6 What About Popper and Kuhn?
The paper aligns its statistical philosophy more closely with Popper’s hypothetico-deductivism than Kuhn’s account, while using limited analogies between Bayesian model checking and scientific anomalies or revolutions.
- Popper: Popper’s falsificationist view emphasizes strong theories whose wide-ranging predictions can be refuted by observations.The authors agree with this general hypothetico-deductive view but argue that Popper’s specific account of statistical testing requires substantial modification.
- Kuhn: Kuhn describes normal science as work within a paradigm whose central presuppositions remain largely unquestioned despite accumulating anomalies.Anomalies may eventually contribute to a crisis and the formation of a new paradigm, though the authors note that sharp historical breaks are difficult to sustain.
- Bayesian practice and Kuhn: Bayesian updating resembles normal science because conditioning on new data treats a model’s prior and likelihood assumptions as fixed, whereas model checking can prompt model replacement or expansion.The authors compare statistical anomalies and model reformulation with Kuhnian crises while cautioning that the analogy should not be pushed too far.
- Scientific change: The authors’ broader view allows scientific progress to combine routine deductive problem solving with occasional larger-scale revisions.They describe this as a weak self-similarity across scales, from local consulting work to broader scientific change.
- Conclusion: The paper concludes that deductive model checking drives statistical progress and motivates inference within complex models known in advance to be false.This position is presented as fundamentally closer to Popper than Kuhn, despite recognizing useful analogies with Kuhnian scientific practice.
7 Why does this matter?
The paper argues that an inductivist philosophy can encourage uncritical reliance on existing Bayesian models. It instead presents model checking and continuous expansion as central to responsible Bayesian data analysis.
- Why philosophy matters: The authors argue that treating Bayesian inference as inductive learning can encourage researchers to fit and compare models without checking them.They regard this as a harmful consequence of the philosophy, not as evidence that Bayesian inference itself is useless.
- Model checking: Posterior predictive checks diagnose model breakdown by simulating data from the model and comparing the simulations with the actual observations.The comparison can often be conducted visually, providing a practical way to identify limitations.
- Limits of model selection: An inductive framework assumes that the true model, or a suitable model among those considered, is already included in the candidate set.The authors say this conflicts with experience in which model failure requires expanding beyond the existing class of models.
- Methodological implications: The proposed workflow combines large models for diverse data, Bayesian parameter-uncertainty summaries, graphical checks, and continuous model expansion.The authors distinguish this practice from model selection or discrete model averaging and present it as compatible with hypothetico-deductivism.
- Practical consequence: Responsible Bayesian practice requires complex models to be checked and falsified rather than treated complacently as the final alternatives.The authors frame learning from model mistakes as the practical consequence of combining powerful inference with model criticism.