Source-linked AI summary

A Unifying Perspective on Probabilities as Model Predictions

Benedikt Höltgen

arXiv:2609.09855v1stat.MLcs.CYcs.LGmath.ST

TL;DR

Probabilistic statements have disputed meanings, and it is unclear when acting on them produces desirable outcomes. The paper treats every probability as the output of a model-dependent prediction method, then uses finite calibration to connect predictions with utility distributions and decision-making. It further argues that calibration can be feasible through induction and the probability calculus, while identifying strong calibration requirements in some settings.

  • Problem

    Probabilistic statements lack a settled interpretation, despite their pervasive role in reasoning and decision-making.

  • Method

    The paper analyzes probabilities as outputs of prediction methods that construct abstractions and transform them into predictions.

  • Results

    Finite calibration can support predicting a policy’s utility distribution and justify expected-utility maximization when cumulative utility is the objective.

  • Takeaways & Limitations

    Calibration can be supported by sample-to-population induction and by probability-calculus constraints on predictors.

  • Takeaways & Limitations

    Exact calibration on all sets of equal utility is strong, though approximate calibration or approximately equal utilities can suffice for approximately correct cumulative-utility predictions.

Abstract

from arXiv · show

Although probabilistic statements are ubiquitous, foundational disagreements persist about their understanding, as exemplified by debates between Bayesians and frequentists; moreover, it is unclear when and why acting on them actually leads to desirable outcomes. Here, we argue that every probability is the output of a \emph{prediction method}, that is, it depends on both a particular way of constructing abstractions and a way of transforming them into predictions. Through this, we provide a unifying perspective on supposedly different kinds of probabilities and show that even supposedly objective ones are model-dependent. We demonstrate that when a finite calibration criterion is met, one can anticipate the distribution of utilities for a given policy and inform successful decision-making on finite sets of events. Based on the notion of prediction methods, inductive arguments, and the probability calculus, we explain the feasibility of the calibration criterion in many settings. Overall, we develop a coherent perspective on probabilities and their use, connecting key intuitions behind other interpretations along the way.

1 Introduction

Probabilistic statements guide everyday and scientific decisions, yet their meaning remains contested among belief-based, frequency-based, and physical interpretations. The paper proposes treating all probabilities as constructed, model-dependent predictions that are useful when finitely calibrated.

  • Probabilistic reasoning shapes ordinary and scientific decisions, from carrying an umbrella to choosing surgery.The paper uses these examples to motivate why understanding probability matters.
  • Probability remains conceptually disputed, including whether it expresses degrees of belief, repeated-trial frequencies, or physical properties.
  • A common framework distinguishes aleatory probabilities grounded in observable frequencies from epistemic probabilities linked to uncertain predictions and credences.
  • The paper argues that all probabilities are constructed and model-dependent rather than belonging to fundamentally separate kinds.
  • Its descriptive account connects probability to finite calibration, while leaving open which predictions normatively deserve the label probability.The perspective also engages common ideas about rational belief and expected-utility decision-making.

2 Behind Every Probability is a Prediction Method

The paper defines probabilities as outputs of prediction methods that construct abstractions and apply predictors to events. This framework covers forecasts, gambling probabilities, relative frequencies, and medical-risk predictions, including cases often regarded as objective.

  • 2.1 Predictors and prediction methods: A prediction method combines a selected predictor with an abstraction of the situation to produce a prediction for an event.The predictor is a function, while the method describes how an actual situation is represented before applying it.
  • 2.1 Predictors and prediction methods: A predictor is defined as a function p : X × A →R over an abstraction space X and an event algebra A.
  • 2.1 Predictors and prediction methods: Rain forecasting illustrates the framework: measurements such as temperature and air pressure form an abstraction that a computer model transforms into a rain prediction.
  • 2.1 Predictors and prediction methods: Event specification and abstraction both involve modelling choices, including which information to use and how events such as rain are defined or measured.
  • 2.2 Supposedly objective probabilities: Symmetry-based gambling probabilities arise from abstractions of fair dice and combinatorial models, even when they appear objective.The framework permits different choices about whether dice properties are inputs, parameters, or built into the predictor.
  • 2.3 Relative frequencies: Relative-frequency methods extend the framework to gambling and medical risks by using observations from similar events or people.

3 Finite Calibration Makes Probabilities Useful

Finite calibration makes probability predictions useful for forecasting event counts and cumulative utility, and for choosing policies that deliver desirable outcomes under specified utility conditions. The paper also shows that calibration alone is insufficient when decisions depend on richer information or unequal utilities.

  • 3.1 Predicting numbers of events: Calibration means that, across a finite event set, the sum of predictions indicates how many events will occur.This provides the basic bridge from probabilistic predictions to anticipated outcomes.
  • 3.1 Predicting numbers of events: In gambling settings, calibrated predictions of repeatable events with fixed payoffs indicate how much to bet for positive long-run outcomes.Calibration predicts event counts, not which particular events occur.
  • 3.2 Predicting cumulative utility: Calibration on sets of equal utility lets predicted utilities determine cumulative utility when utilities are assigned to binary outcomes.The paper formalizes this by requiring calibration on sets sharing each possible utility value.
  • 3.4 Calibration revisited: Exact calibration assumptions are strong, while approximate calibration can still support approximately correct cumulative-utility predictions.The paper also notes that calibration must be supplemented by sharpness because more informative predictions can be harder to calibrate.
  • 3.2 Predicting cumulative utility: Expected-utility maximisation actually maximises cumulative utility when predictions are calibrated on equal-utility sets and cumulative utility is the objective.Under these conditions, the policy with higher expected cumulative utility provides higher cumulative utility.
  • 3.3 Probability calibration: In the umbrella example, the utility-maximising threshold is pi > 0.25, and calibration makes that threshold optimal among threshold-based policies.The threshold follows from comparing the predicted utilities of bringing versus not bringing an umbrella.

4 How Is Calibration Possible?

Calibration can be feasible when predictions are generated by suitable methods and evaluated on related finite sets. Inductive and probabilistic arguments support transferring calibration, while the paper retains explicit scope conditions and practical limitations.

  • Calibration cannot be assumed for arbitrary prediction sets, so understanding the methods that generate predictions is essential.The paper distinguishes the calibration criterion from the conditions under which it can realistically be achieved.
  • Prediction methods can be designed to be approximately calibrated on sets of interest, meaning they neither systematically over-predict nor under-predict.Examples include rain forecasts and machine-learning models calibrated on sets of equal prediction.
  • Proposition 8 uses concentration for samples drawn without replacement to show that most sufficiently large samples have calibration error close to the population error.The result is obtained by applying Hoeffding’s inequality to finite populations.
  • An average population calibration error of 0.2 yields calibration error below 0.05 for fewer than 10% of samples of size 200 and 0.4% of samples of size 500.These are upper-bound illustrations; the actual fraction of non-representative samples may be lower.
  • Past calibration indicates future calibration on similar sets when the sample can reasonably be treated as unbiased, but this does not prove induction.The relevant similarity between past and future sets remains situation-dependent.
  • The same population argument extends to mixtures of calibrated prediction methods, although future calibration can never be guaranteed.The paper also derives calibrated predictions from related calibrated predictions through probability-calculus properties such as non-negativity, normalisation, and additivity.
  • A single violation of the probability axioms can make a best-response policy sub-optimal for some utility functions despite calibration on another forecast.The paper therefore presents the probability calculus as a sound and complete way to generate calibrated predictions on certain related sets, without claiming that every predictor must obey it.

5 A Unified Interpretation of Probability

The paper treats probabilities as constructed outputs of prediction methods rather than intrinsic properties or purely subjective beliefs. Finite calibration connects these predictions to policy evaluation and decision-making while leaving the account agnostic about determinism.

  • The framework remains agnostic about determinism because inaccessible quantum-mechanical probabilities may differ from useful predictions available in practice.The authors argue that quantum mechanics does not establish that everyday probabilities should be identified with fundamental physical probabilities.
  • Prediction methods can operate mechanically and need not correspond to an entity possessing degrees of belief.The paper nevertheless allows probabilities to model human or machine decisions as if they involved beliefs.
  • The account explains links to relative frequencies while rejecting finitary frequentism as too narrow an evaluation criterion.Calibration restricted to equal-probability sets becomes equivalent to a finitary frequentist definition, but the authors treat that as a special case.
  • Probabilities depend on both an abstraction of a situation and a predictor, so they are model-dependent rather than properties of events.The same event can receive different probabilities from different choices of abstraction and prediction method.
  • Finite calibration can justify policies by allowing predicted utilities to track cumulative utilities across relevant event sets.Expected utility maximization is supported when calibration holds on sets of equal utility.

6 Comparison with Conventional Interpretations

The paper compares its prediction-method account with conventional interpretations and argues that it combines empirical justification with constructed, model-dependent probabilities. It retains useful insights from Bayesian, frequentist, and related views without treating their probabilities as uniquely true.

  • The framework differs from best-systems and logical accounts because its predictors are functions of abstractions and events, not objective relations between propositions.The comparison preserves similarities while rejecting uniquely objective probabilistic laws or relations.
  • Unlike Bayesianism, it grounds probabilities in prediction methods rather than degrees of belief while retaining the idea that probabilities are constructed.Prediction methods aim to track structure in sets of observations.
  • Hypothetical frequentism shares the focus on repeated-trial ratios, but the paper emphasizes dependence on reference classes and conceptualizations.Different reference classes can yield different probabilities for an individual event.
  • The account treats probabilities as constructed rather than discovered while grounding their justification in empirical observations.Calibration connects empirical evaluation with successful decision-making and induction.
  • The paper presents calibration as a central bridge between empirical claims, probability assignments, and successful action.This connection is offered as a way to avoid both neglect of empirical claims and unnecessary objective probability stipulations.

7 Conclusion

The conclusion presents prediction methods and finite calibration as a pragmatic framework unifying probability, abstraction, forecasting, and decision-making. It also cautions against treating machine-learning outputs as samples from a context-independent true distribution.

  • Finite calibration can make predicted utility distributions informative for evaluating policies and can support expected utility maximization.The predicted sum of utilities can match actual cumulative utility under the criterion.
  • Prediction methods unify rain forecasts, gambling odds, and other probabilities by combining abstractions, forecasting, and empirical evaluation.The framework’s novelty lies especially in connecting these elements to successful decision-making.
  • The account motivates evaluating calibration beyond sets of equal predictions in machine learning.It also warns against invoking a true distribution from which one can sample outside highly controlled settings.
  • The perspective extends to causal inference by highlighting calibration as relevant to probabilistic reasoning beyond forecasting.The paper identifies causal inference as an area addressed in concurrent work from this perspective.

A.1 Approximate calibration

The appendix extends cumulative-utility guarantees from exact to approximate calibration and notes that alternative error measures may be useful when over- and under-prediction have different value.

  • Approximate calibration bounds the difference between predicted and cumulative utility by a chosen tolerance ϵ.The bound is formulated for calibration error on each set of equal utility.
  • The approximate-calibration proposition assumes positive utilities, with shifting available when utilities are otherwise nonpositive.The result replaces the exact calibration condition under this positivity assumption.
  • The analysis uses symmetric ℓ1 loss but identifies asymmetric error functions as a possible alternative.Asymmetric losses may better represent settings where under-prediction and over-prediction are valued differently.

A.2 Approximate utility level sets

Approximate utility calibration replaces exact utility-level matching with bins of width δ, yielding a bounded discrepancy between predicted and cumulative utility.

  • A.2 Approximate utility level sets: The resulting bound on cumulative-utility mismatch depends on δ and the number of predictions d.This quantifies the cost of replacing exact utility-level calibration with approximate level sets.
  • A.2 Approximate utility level sets: Utilities are grouped into bins of size at most δ to relax calibration requirements across approximately equal utility levels.The relevant utility range is partitioned from the minimum to maximum utility values.
  • A.2 Approximate utility level sets: The maximal mismatch is attained when prediction errors of opposite signs pair with utilities at opposite edges of each bin.The construction places positive and negative discrepancies at utilities separated around each bin midpoint.

A.3 Imprecise calibration

Imprecise calibration extends calibration from point predictions to intervals, allowing cumulative utility to be predicted as a range while accommodating risk-sensitive decisions.

  • A.3 Imprecise calibration: Imprecise forecasts represent each prediction as an interval with lower and upper probabilities.The interval is encoded as a tuple containing its lower and upper probability.
  • A.3 Imprecise calibration: Imprecise calibration requires observed event counts to lie between the sums of lower and upper predictions.This weakens exact matching between event frequencies and summed point predictions.
  • A.3 Imprecise calibration: The vacuous forecast that always predicts (0, 1) is always imprecisely calibrated.The criterion therefore permits forecasts that provide the full binary-probability range.
  • A.3 Imprecise calibration: Imprecise calibration can be applied to precise predictions by widening each point prediction to [p_i − ε, p_i + ε].The widening parameter ε may decrease with d to reflect lower variance on larger sets.
  • A.3 Imprecise calibration: Under imprecise calibration on equal-utility sets, cumulative utility can be predicted as an interval rather than a single value.This supports selecting policies using lower or higher utility estimates to reflect risk aversion or risk seeking.

B Conditional probabilities, generalised

The generalized conditional-probability analysis characterizes when predictions for A|B are calibrated on the subset of cases where B occurs.

  • B Conditional probabilities, generalised: A predictor calibrated for A_i ∩ B and B, with equal predictions for B, can be analyzed for conditional predictions A_i|B.The characterization relates calibration on joint events and conditioning events to calibration on the restricted cases.
  • B Conditional probabilities, generalised: The generalized result provides an equivalence and a sufficient condition for calibration on the cases where B occurs.The supplied derivation states these conditions through the displayed characterization and its sufficient condition.
  • B Conditional probabilities, generalised: When all events are A and all inputs are x, the familiar definition of conditional probability is necessary and sufficient for calibration of A|B.This applies to predictions based on x on the subset of steps where B occurs.
  • B Conditional probabilities, generalised: Predicting A|B can achieve calibration on cases where B occurs even when predictions for A ordinarily would not.Conditioning therefore changes the prediction target to match the restricted calibration set.
Loading 2609.09855v1…