Source-linked AI summary
A review of predictive uncertainty estimation with machine learning
Hristos Tyralis, Georgia Papacharalampous
TL;DR
Probabilistic prediction and forecasting can communicate more information than point predictions, but related concepts and methods have not been structured under a holistic view. This review synthesizes predictive uncertainty estimation, assessment metrics, and developments from early statistical models to recent machine-learning algorithms, while classifying the field and discussing emerging challenges.
Problem
Related concepts and methods for probabilistic prediction and forecasting have not been formalized and structured under a holistic view of the field.
Method
The paper reviews probabilistic forecasting and prediction with machine-learning algorithms, including predictive uncertainty estimation and scoring functions or rules.
Results
The review synthesizes seemingly unrelated methodologies and classifies material spanning early statistical models and recent machine-learning algorithms.
Takeaways & Limitations
The review supports understanding how to develop new algorithms tailored to users’ needs by relating recent advances to fundamental concepts.
Takeaways & Limitations
Scoring rules may not meet the requirements of all probabilistic prediction tasks.
Abstract
from arXiv · showhide
Predictions and forecasts of machine learning models should take the form of probability distributions, aiming to increase the quantity of information communicated to end users. Although applications of probabilistic prediction and forecasting with machine learning models in academia and industry are becoming more frequent, related concepts and methods have not been formalized and structured under a holistic view of the entire field. Here, we review the topic of predictive uncertainty estimation with machine learning algorithms, as well as the related metrics (consistent scoring functions and proper scoring rules) for assessing probabilistic predictions. The review covers a time period spanning from the introduction of early statistical (linear regression and time series models, based on Bayesian statistics or quantile regression) to recent machine learning algorithms (including generalized additive models for location, scale and shape, random forests, boosting and deep learning algorithms) that are more flexible by nature. The review of the progress in the field, expedites our understanding on how to develop new algorithms tailored to users' needs, since the latest advancements are based on some fundamental concepts applied to more complex algorithms. We conclude by classifying the material and discussing challenges that are becoming a hot topic of research.
1. Introduction
The introduction motivates predictive uncertainty estimation as a way to communicate more information than point predictions and support better-informed decisions. It presents a unified review of concepts, assessment metrics, statistical and machine-learning methods, special applications, and emerging challenges.
- Motivation: Probability distributions communicate more information than single-value point predictions and characterize uncertainty about future outcomes.The introduction distinguishes predictive distributions from point predictions and links them to uncertainty estimation.
- Foundations: Bayesian predictive uncertainty estimation updates parameter distributions conditional on data and integrates the posterior to obtain predictive uncertainty.The approach is described as an early formal strategy for estimating prediction distributions.
- Assessment and estimation: Quantile loss and proper scoring rules provide decision-theoretic tools for estimating and assessing probabilistic predictions.Quantile loss can estimate predictive-distribution quantiles when minimized in regression, while scoring rules assess probabilistic predictions.
- Review scope: The review synthesizes statistical and machine-learning models, including generalized distributional models, random forests, boosting, and deep learning algorithms.It also connects simpler and more complex models through their underlying loss functions and probabilistic-prediction components.
- Prediction and forecasting: Ideas from probabilistic prediction and forecasting can transfer directly between fields, including when IID machine-learning algorithms are applied to forecasting.Such algorithms can frequently show good performance despite not exploiting temporal dependence information.
- Applications and challenges: The review examines special cases such as time series, spatial and spatiotemporal prediction, extreme events, measurement errors, and combinations of techniques.It concludes by synthesizing results, outlining a future outlook, and discussing challenges in probabilistic prediction.
2. Definitions and history
The section defines predictive distributions for future responses and traces Bayesian modelling from parameter uncertainty to predictive inference. It also positions the review as a holistic account of probabilistic prediction, forecasting, algorithms, and assessment.
- Definitions: The prediction problem is to estimate the distribution of future responses given a statistical model, observed data, and future predictor values.This distribution is called the predictive distribution.
- Bayesian statistical modelling: Integrating the parameter posterior produces the predictive distribution, whereas a point estimate of parameters ignores estimation uncertainty.The review presents Bayesian modelling as a way to treat uncertain parameters as random variables.
- Bayesian statistical modelling: Bayesian regression combines a likelihood-based statistical model with a prior to estimate the posterior distribution of model parameters.Observed realizations of the response and predictors are used to estimate p(θ|y, x).
- Forecasting: The independence assumption underlying the basic predictive-distribution derivation is unsuitable for time series, where temporal dependence can lead to inferior predictive performance.The paper states that this case is reformulated later for forecasting applications.
- Review scope: Existing reviews cover narrow algorithm classes or scoring, but generally do not integrate probabilistic-prediction algorithms and their assessment across the field.The review aims to present a complete view of topics and their interplay, including loss functions and machine learning methods.
3. Loss functions for assessing probabilistic predictions
The section develops scoring functions for probabilistic predictions, emphasizing consistency with target functionals such as quantiles and expectiles. Quantile scores use asymmetric penalties, while expectile scores provide a least-squares analogue with related but distinct interpretation.
- Scoring-function foundations: Scoring functions and scoring rules provide the theoretical basis for assessing probabilistic predictions and have shaped machine learning loss-function development.The review grounds its treatment in scoring-function theory for probabilistic predictions.
- Quantile predictions: Quantile loss asymmetrically weights errors by their sign, enabling predictions at levels other than the median and targeting the correct proportion of observations below the prediction.At level 1/2, the loss becomes symmetric.
- Quantile predictions: The quantile loss function is strictly consistent for quantiles of distributions with finite first moments, so expected-score optimization identifies the target quantile.Generalized piecewise linear scores characterize consistent quantile scoring functions.
- Expectile predictions: Expectile loss is an asymmetric squared-error criterion, reducing to half the squared-error function at τ = ½ and yielding strictly consistent scores under strict convexity.Expectiles are described as least-squares analogues of quantiles.
- Expectile predictions: Expectiles can be coherent and informative about loss magnitude, whereas quantiles describe loss frequency; expectiles may nevertheless be used less often because of interpretability concerns.The paper notes that quantiles are not coherent risk measures and that expectiles are less frequently used.
3.3 Proper scoring rules
Proper scoring rules assess probabilistic predictions by comparing issued distributions with outcomes, with strict propriety rewarding the true predictive distribution. Their decompositions distinguish uncertainty, resolution or sharpness, and reliability, while specialized rules support quantiles, intervals, and ensemble forecasts.
- A proper scoring rule optimizes expected score when the issued predictive distribution matches the true predictive distribution; strict propriety makes equality unique.
- A scoring rule assigns a score to a predictive distribution and an observed outcome, extending point-error measures by using a distribution as its first argument.
- The expected score decomposes into uncertainty, resolution or sharpness, and reliability, separating climatological uncertainty from forecast informativeness and calibration.
- Continuous Ranked Probability Score (CRPS): The CRPS evaluates full predictive distributions through their cumulative distribution functions and can assess sample-based forecasts such as simulations.
- Continuous Ranked Probability Score (CRPS): For ordered ensembles, CRPS can be estimated from ensemble members, whereas the corresponding expression for unordered ensembles is not recommended.
- Scoring rules for quantiles and intervals: Quantile and interval scoring rules evaluate predictions at specified quantile levels, including central intervals whose coverage is determined by paired lower and upper quantiles.
- Scoring rules for quantiles and intervals: Comparing prediction intervals at a specified coverage probability without specified quantiles is described as difficult and perhaps unsolvable.
3.4 Some more proper scoring rules for assessing probabilistic predictions
The review surveys additional proper scoring rules and related validation concepts for probabilistic predictions. It distinguishes local density-based scoring, skill scores, minimax scoring, and backtesting, while emphasizing that scoring rules should match users' requirements.
- The log score is a local proper scoring rule connected with likelihood inference, but unlike CRPS it is restricted to predictive densities.
- Other scoring rules target moments, simultaneous events, temporal dependence, joint predictions, asymmetric errors, survival outcomes, and distributional distances.
- Skill scores standardize scoring rules for comparisons across multiple cases, but they are not necessarily proper.
- Scoring rules can be selected or constructed to meet prediction users' specified requirements.
- Liu et al. propose minimizing a scoring rule over all possible true distributions, converting the problem into a minimax formulation because the true data distribution cannot be learned by a model.
- Backtesting validates a single model, whereas elicitability concerns comparison of multiple models; VaR backtesting examines nominal conditional coverage and violations.
4. Early history and simple models
The review introduces simpler probabilistic models before machine-learning models because their theory underlies more complex approaches. It covers Bayesian linear regression, quantile regression, time-series models, and copula models.
- Simpler probabilistic models are presented first to clarify how more complex and accurate machine-learning models can be built.
- The early models include Bayesian statistical ordinary linear regression, linear-in-parameters quantile regression, time-series models, and copula models.
- The review states that machine-learning theory builds on simpler models by adding complexity, resulting in improved performance.
4.1 Bayesian statistical models
Bayesian statistical models provide predictive distributions through analytical results in Gaussian regression and simulation-based methods when explicit forms are unavailable. Their practical use is constrained by data-generating-process assumptions and increasing computational costs.
- Gaussian regression: Bayesian linear regression yields a predictive distribution with a multivariate Student form and n − p degrees of freedom.Its location and scale depend on future predictors, the predictor matrix, responses, and an estimated squared scale.
- Simulation-based prediction: Predictive distributions for Bayesian statistical models can be estimated with MCMC simulations.Simulation may be appropriate, and perhaps the only means to proceed, beyond the Gaussian case when predictive distributions are not explicit.
- Validation and assumptions: Bayesian modelling is mainly concerned with in-sample performance and requires correct specification of the data-generating process.Cross-validation using out-of-sample performance and various loss functions has also been studied.
- Validation and assumptions: As datasets increase, the computational cost of cross-validation may increase significantly.The passage notes that Bayesian statistical models are mostly applied to small datasets.
- Approximate Bayesian methods: Approximate Bayesian Computation addresses increasing dataset size and likelihood intractability by retaining simulated data close to observations under a distance metric.It simulates artificial data from prior-distributed parameters and has been reported to provide efficient predictions with nearly identical, although inferior, results to exact ones.
- Approximate Bayesian methods: Focused Bayesian prediction does not require correct specification of the data-generating process and updates predictions using a specified measure such as a proper scoring rule.The method is also applicable to improving combinations of algorithms.
4.2 Quantile and expectile regression linear in parameters
Quantile and expectile regression extend probabilistic prediction beyond the mean through loss-based estimation of distributional functionals. The section surveys regularization, Bayesian formulations, multivariate extensions, and practical problems including crossing and nondifferentiable optimization.
- Quantile and expectile losses: Quantile predictions can be assessed using quantile loss, and linear-model parameters can be estimated by minimizing expectile loss.Expectile regression has similar properties to quantile regression but differs because of the respective loss functions.
- Quantile regression: Quantile regression is appropriate for predicting functionals beyond the mean, known predictive distributions, outliers, and heteroscedasticity.The passages identify robustness to outliers and applicability when variance depends on covariates.
- Quantile regression: Regularization, including ridge regression and lasso, improves inference or prediction when parameter estimation is unstable or models are high-dimensional.Regularization has been implemented in quantile regression models.
- Bayesian quantile regression: Bayesian quantile regression is useful when parameter inference is required or optimization is difficult, but parametric likelihoods are unavailable and working likelihoods are usually used.This limitation distinguishes the Bayesian treatment from standard parametric likelihood specification.
- Extensions: Multivariate quantile regression predicts quantiles of d ≥ 2 variables simultaneously, increasing the problem’s complexity.Directional approaches reduce the multivariate problem to univariate marginal problems, while direct approaches include spatial, elliptical, and depth-based quantiles.
- Practical problems: Quantile crossing occurs when independently estimated conditional quantiles violate their expected ordering across quantile levels.Remedies include sorting estimated quantiles and enforcing non-crossing constraints.
- Practical problems: Because quantile loss is not everywhere differentiable, gradient-based optimization methods are not always applicable.Approximations to the quantile loss function have been employed for efficient optimization.
- Quantile and expectile regression: Expectile regression has developed more slowly than quantile regression, possibly because expectiles are difficult to interpret.The passage attributes this possibility to interpretability issues.
4.3 Forecasting with time series models
Time-series probabilistic forecasting conditions future distributions on observed history and must account for temporal dependence. The review covers Bayesian, approximate Bayesian, bootstrap, quantile, expectile, and approximation-based approaches across several time-series model families.
- Forecasting setup: Probabilistic forecasting targets the distribution of a future variable y_t+h conditional on observations through time t.One-step forecasting uses h = 1, whereas multi-step forecasting uses h > 1.
- Bayesian forecasting: Temporal dependence complicates Bayesian forecasting because the likelihood and predictive calculation must incorporate the stochastic process structure.The one-step-ahead predictive distribution integrates the conditional future distribution over the posterior distribution of parameters.
- Computational considerations: The length of a time series may prohibit Bayesian methods from being practically applicable.Approximate Bayesian computation is presented as a possible solution with little cost in predictive performance.
- Alternative forecasting methods: Bootstrap forecasting approximates the innovation distribution by resampling residuals from a fitted time-series model.Innovations are described as the random component that models the model error.
- Quantile and expectile forecasting: Quantile-based time-series models include CAViaR models, which specify quantile evolution over time using an autoregressive process.CARE models are similar but are estimated with an expectile loss function.
- Alternative forecasting methods: Alternative probabilistic forecasting techniques include Gaussian analytical, variance, Taylor, asymptotic, and error-distribution approximations.The review notes that relevant ideas may also apply to general machine-learning frameworks.
4.4 Copula-based regression models
Copula-based regression models construct conditional distributions by separating marginal distributions from dependence modelling. The review covers semiparametric estimation, quantile and survival extensions, high-dimensional regularization, and specialized copula classes.
- Copula foundations: Copulas model multivariate distributions with uniform marginal distributions and separate dependence modelling from marginal distributions.This separation facilitates the modelling procedure.
- Copula-based regression: Copula-based regression substitutes predictor-variable marginal distributions into a copula to obtain a conditional distribution G given x and a specified dependence model C.The approach extends naturally to probabilistic regression predictions.
- Model assumptions: Copula-based regression is parametric in the copula function, making correct copula specification important.When the copula is misspecified, results may be largely unreliable.
- Estimation procedures: A semiparametric estimator can model the copula parametrically while modelling marginal distributions non-parametrically.This estimation strategy is attributed to Noh et al. (2013).
- Extensions: Copula-based regression has been extended to quantile regression for IID data and time series, and to quantile estimation in survival analysis.Vine copulas have also been used for direct conditional quantile estimation.
- Extensions: High-dimensional Gaussian copula regression has been combined with variable-selection methods that transform selection into a multiple-testing problem.One reported purpose is reducing redundant variables due to the stopping rule.
- Extensions: Vine copulas have been applied to mixed data, conditional heteroscedasticity, large-scale, nonlinear, and non-Gaussian data.Other surveyed approaches include Gaussian copula regression and Bayesian copula-based regression with nonlinear predictors.
5. Machine learning algorithms
The review shows how established probabilistic-prediction concepts can be transferred to flexible machine-learning models, including GAMLSS, random forests, and boosting.
- Machine-learning models can issue probabilistic predictions by combining them with fundamental concepts introduced earlier.
- GAMLSS: GAMLSS models directly estimate predictive-distribution parameters, allowing the full distribution to be modelled rather than only its functionals.They can model parameters beyond mean and variance and accommodate any probability distribution.
- GAMLSS: GAMLSS uses additive functions, random effects, and penalized-likelihood optimization, but requires specifying the distribution, link function, and estimation procedure.
- Random forests: Random-forest variants estimate conditional distributions or quantiles using tree responses and indicator functions instead of ensemble averages.Quantile regression forests are consistent for quantile estimation under certain general conditions.
- Random forests: Variants modify random forests through bias correction, quantile-loss or gradient-based splitting, local-linear adjustments, and distribution-specific likelihoods.
- Boosting: Boosting trains base learners sequentially to minimize a loss function, adding learners to correct errors made by earlier learners.
6. Neural networks and deep learning
Neural networks and deep-learning models support probabilistic prediction through Bayesian approximations, ensembles, dropout, quantile methods, and distributional regression.
- Bayesian methods: Bayesian neural networks estimate posterior and predictive distributions for model parameters and outputs.Variational inference approximates computationally expensive posterior distributions by minimizing KL divergence.
- Monte Carlo methods: Dropout approximates probabilistic deep Gaussian processes in a Bayesian context while being computationally faster.Predictive uncertainty can be estimated from the variance of the resulting neural-network ensemble.
- Monte Carlo methods: Neural-network ensembles can approximate Bayesian uncertainty through batch normalization, randomized architectures, or marginalization over network depth.
- Quantile regression: Quantile-regression neural networks estimate multiple quantiles, and joint estimation can mitigate quantile crossing by sharing information across levels.An L1 penalty on the loss gradient is one proposed way to address crossing quantiles.
- Distributional regression: Distributional regression with neural networks estimates parameters of target distributions and can represent richer, non-Gaussian families.
- Other approaches: Other probabilistic deep-learning approaches include conformal prediction, transformation models, generative adversarial networks, and scoring-function-based training.
7. A representative simple example
A simulation compares quantile regression and GAMLSS for probabilistic prediction under known Gaussian data-generating conditions. Both improve with larger samples, while GAMLSS remains better in this setting.
- Simulation design: The experiment simulates y = x1 + x2 + ε with Gaussian predictors and noise, using separate training and testing samples.The main simulation uses 100 samples, split into 50 for training and 50 for testing.
- Evaluation: Quantile regression and Gaussian GAMLSS produce probabilistic predictions at multiple quantile levels, evaluated by coverage probabilities and quantile scores.
- Results: Both methods become more accurate as sample size increases, and their performance difference decreases.
- Results: 0.07344 and 0.05211 are the mean coverage probabilities at α = 0.05 for sample sizes 100 and 1,000, respectively, for GAMLSS.
- Interpretation: GAMLSS performs better because the simulation specifies the dependent variable’s distribution, giving the parametric model an advantage.
- Interpretation: Non-parametric models improve at a higher rate as sample size increases but remain worse than parametric models in this simulation.The situation could reverse with real data when the probability distribution is misspecified.
8. Combinations of algorithms
Combining probabilistic predictions can improve forecasts through weighted averaging, quantile or distributional pools, Bayesian model averaging, and stacking.
- Overview: Ensemble combinations can improve probabilistic predictions, extending the benefits of combining algorithms beyond point-prediction performance.
- Weighted averaging: Weighted averaging combines predicted distributions, densities, quantiles, or other distributional functionals.
- Weighted averaging: The equally weighted linear pool is difficult to beat in practice, and its attempts are more successful as data size increases.
- Weighted averaging: Averaging quantiles can produce sharper forecasts than averaging distributions, while equal weighting performs no worse than the average score of individual algorithms.
- Bayesian model averaging: Bayesian model averaging combines model-specific predictive distributions using posterior model probabilities and can improve underdispersive forecast ensembles.
- Stacking: Stacking estimates combination weights from out-of-sample performance using consistent scoring functions or proper scoring rules.It is favored over BMA when the candidate algorithms do not include the true data-generating model.
9. Special cases
Special cases require probabilistic prediction methods that exploit temporal or spatial dependence, address extremes and measurement errors, or combine flexible algorithms with distributional modeling. The review surveys specialized statistical, machine-learning, deep-learning, and hybrid approaches for these settings.
- Temporal and spatial prediction: Temporal and spatial dependence can be exploited to improve predictive performance, motivating specialized models for time series, spatial, and spatio-temporal data.Time-series models and specialized deep-learning architectures target temporal dependence, while spatial models address dependence and irregular measurement locations.
- Forecasting methods: Hybrid forecasting models combine complementary algorithms, such as neural networks with GARCH variance modeling, spline quantile functions, or exponential smoothing.The reviewed combinations respectively model means and variances, represent response distributions, or deseasonalize and normalize series.
- Extremes and other special cases: The review also covers quantile-regression neural networks, Monte Carlo and bootstrap methods, distributional regression, Gaussian processes, Markov random fields, and other specialized approaches.These methods include dropout or bootstrap prediction intervals, neural estimation of bounded-distribution parameters, and spatial models for regular grids.
- Temporal and spatial prediction: Modelling multiple time series simultaneously can exploit additional information and lead to large predictive-performance improvements over modelling each series separately.Examples combine convolutional or autoregressive components with quantile or distributional regression.
- Extremes and other special cases: Specialized probabilistic methods address extremes by modeling distribution tails or extreme quantiles, but distributional-regression parameter estimates may be unstable in such cases.Heavy-tailed distributions are often suitable for variables with substantial tail mass; quantile regression is also discussed for these settings.
10. Summary and future outlook
The review organizes probabilistic prediction around problem type, prediction target, and scoring methodology, showing how foundational concepts and machine-learning algorithms can be combined for task-specific models. It concludes that probabilistic prediction applications are expected to increase in academia and industry while identifying scoring-function choice and extreme-quantile prediction as continuing challenges.
- Model construction: The review presents three foundational areas—Bayesian concepts, predictive-problem definitions, and scoring rules—as components for building models tailored to specific tasks.Scoring rules support fitting parametric models without Bayesian simulations and non-parametric models without distributional misspecification.
- Problem formulation: Probabilistic prediction problems differ by independence, temporal dependence, spatial dependence, and special interests such as measurement errors or extremes.Most machine-learning models address conditionally independent responses, whereas temporal and spatial dependence motivate specialized models.
- Problem formulation: The prediction target may be a functional of the predictive distribution or the full distribution, with the latter providing more information but requiring harder estimation.Quantiles can be fitted with consistent scoring functions, while distributional regression and proper scoring rules target full predictive distributions.
- Model construction: Elements from these foundations and machine-learning algorithms can be combined, for example by using distributional boosting to estimate a full predictive distribution.The example optimizes a proper scoring rule applied to the dependent-variable distribution within a machine-learning algorithm.
- Challenges and outlook: Scoring-function choice matters because different consistent scoring functions for the same elicitable functional can yield different assessments.Under stated assumptions, however, the choice does not affect assessment.
- Challenges and outlook: Applications of probabilistic predictions using machine-learning algorithms are anticipated to increase in both academia and industry.The review also identifies extreme-quantile prediction as a continuing research challenge.