Source-linked AI summary
A General Framework for Updating Belief Distributions
Pier Giovanni Bissiri, Chris Holmes, Stephen Walker
TL;DR
Bayesian inference becomes difficult when the true sampling distribution is unknown, overly complex, or unavailable for the parameter of interest. The paper replaces the likelihood connection with loss functions and decision-theoretic updating. The resulting framework recovers ordinary Bayesian updating under self-information loss and supports coherent inference with more general information.
Problem
Conventional Bayesian inference requires a complete sampling model, which is problematic when the data-generating distribution is unknown or when the parameter of interest does not index a density family.
Method
The framework connects information to parameters through loss functions and defines posterior beliefs using decision-theoretic cumulative-loss updating, including non-stochastic information.
Results
The procedure recovers traditional Bayesian updating when the self-information loss is appropriate and provides coherent updates for more general information and parameter settings.
Takeaways & Limitations
Bayesian-style subjective inference can be extended to targets and information that cannot be handled naturally by a complete likelihood model.
Takeaways & Limitations
Some calibration procedures still require specifying a sampling distribution for x, while the illustrative clustering analysis assumes constant memberships and shared change points over time.
Abstract
from arXiv · showhide
We propose a framework for general Bayesian inference. We argue that a valid update of a prior belief distribution to a posterior can be made for parameters which are connected to observations through a loss function rather than the traditional likelihood function, which is recovered under the special case of using self information loss. Modern application areas make it is increasingly challenging for Bayesians to attempt to model the true data generating mechanism. Moreover, when the object of interest is low dimensional, such as a mean or median, it is cumbersome to have to achieve this via a complete model for the whole data distribution. More importantly, there are settings where the parameter of interest does not directly index a family of density functions and thus the Bayesian approach to learning about such parameters is currently regarded as problematic. Our proposed framework uses loss-functions to connect information in the data to functionals of interest. The updating of beliefs then follows from a decision theoretic approach involving cumulative loss functions. Importantly, the procedure coincides with Bayesian updating when a true likelihood is known, yet provides coherent subjective inference in much more general settings. Connections to other inference frameworks are highlighted.
3.1 Annealing.
The paper develops ways to calibrate the relative influence of data loss and prior information, including annealing, hierarchical, operational, and subjective approaches. These methods extend belief updating beyond settings requiring a complete sampling model.
- Annealing: The weighting parameter w controls the relative influence of data loss and prior information in the update.Values below 1 make data less influential, while values above 1 give the data loss greater prominence than in the Bayesian update.
- Unit information loss: w can be calibrated by matching prior expected loss with expected data loss, although this requires specifying a joint belief distribution for x and θ.An observed unit-information loss provides an alternative calibration based on the data rather than a prior expectation.
- Unit information loss: The unit-information construction requires specifying a sampling distribution for x, so the empirical distribution function can instead be used as an empirical choice.The empirical distribution function substitutes for the unknown F0 in the finite-sample procedure.
- Hierarchical loss: A hierarchical loss treats w as unknown and adds a penalty ξl(w), making the choice of ξ less crucial because it can be absorbed into the prior.For a scale parameter, the paper considers l(w) = log w and allows ξ to be assessed subjectively.
- Operational characteristics and subjective calibration: Operational calibration selects w so posterior quantiles match frequentist confidence-interval coverage at a chosen error level.The paper also relates w to Bayes factors: w scales how loss differences translate into relative evidence between parameter values.
- General forms of information: The framework defines conditional distributions from non-stochastic information by connecting that information to θ through a loss function.This extends the updating approach to settings where the information used for updating is not a stochastic event already included in a probability model.
4.2 Partial information.
The paper illustrates partial-information Bayesian updating by using loss functions that target parameters of interest without requiring a full generative model. Applications to survival analysis and clustering show that the framework can retain relevant information while quantifying uncertainty in the target structure.
- Survival analysis: Partial likelihood is treated as a valid Bayesian update when it represents partial self-information loss.The authors contrast this motivation with conventional Bayesian Cox-model analyses that require specifying the baseline hazard.
- Survival analysis: The general Bayes Factors show strong association evidence around marker 10,000 and broadly agree with conventional likelihood-ratio evidence at strongest markers.The comparison reports greater dispersion for weaker associations, especially when plotted against likelihood-ratio-test p-values.
- Clustering: In clustering, conventional Bayesian analysis must introduce mixture components, nuisance parameters, priors, and a symmetric likelihood although the target is the partition structure.The proposed approach instead places a prior directly on the partition and connects observations to it through a loss function.
- Clustering: For the election data, the loss-based posterior provides strong evidence for clustering across both States and time, with the maximum posterior probability favoring three State groups.The analysis calibrates the loss parameter and uses MCMC to compare configurations with differing numbers of State groups and time-series change points.
- Clustering: Posterior co-clustering probabilities agree with Hartigan’s reported cocluster while quantifying considerable uncertainty in pairing Virginia and North Carolina.The framework also identifies strong evidence that time-series change points occur late in the series.
- Implications: The discussion states that loss functions on probability-measure spaces support coherent updating, recover Bayes’ rule under self-information loss, and accommodate robust-estimation losses and partial information.The framework replaces the restrictive probability-model connection with a loss-function connection to the parameter of interest.