Source-linked AI summary
Forecaster's Dilemma: Extreme Events and Forecast Evaluation
Sebastian Lerch, Thordis L. Thorarinsdottir, Francesco Ravazzolo, Tilmann Gneiting
TL;DR
The paper examines how evaluating forecasts only when extreme outcomes occur can distort assessments and discredit skillful forecasts, especially with low signal-to-noise ratios. It combines theoretical analysis, simulations, and a real-data study of U.S. inflation and GDP growth to study the dilemma and potential remedies. The paper finds that all-case evaluation is important, while weighted scoring rules offer coherent emphasis on extremes but limited power benefits over unweighted scores.
Problem
Forecast evaluation often conditions on extreme observed outcomes, despite theoretical assumptions that apply to the full joint distribution of forecasts and observations.
Method
The paper uses theoretical arguments, simulation experiments, and a real-data study of probabilistic U.S. inflation and GDP growth forecasts.
Results
Skillful forecasts can be discredited when signal-to-noise ratios are low, while proper weighted scoring rules provide coherent emphasis on extremes but generally limited power gains over unweighted scores.
Takeaways & Limitations
Forecast evaluation should consider all available cases, and weighted scoring rules can emphasize regions of interest while preserving decision-theoretic coherence.
Takeaways & Limitations
The paper uses a quadratic sample-based approximation to the logarithmic score in its Bayesian analysis, leaving more efficient and theoretically principled approximations for future work.
Abstract
from arXiv · showhide
In public discussions of the quality of forecasts, attention typically focuses on the predictive performance in cases of extreme events. However, the restriction of conventional forecast evaluation methods to subsets of extreme observations has unexpected and undesired effects, and is bound to discredit skillful forecasts when the signal-to-noise ratio in the data generating process is low. Conditioning on outcomes is incompatible with the theoretical assumptions of established forecast evaluation methods, thereby confronting forecasters with what we refer to as the forecaster's dilemma. For probabilistic forecasts, proper weighted scoring rules have been proposed as decision theoretically justifiable alternatives for forecast evaluation with an emphasis on extreme events. Using theoretical arguments, simulation experiments, and a real data study on probabilistic forecasts of U.S. inflation and gross domestic product growth, we illustrate and discuss the forecaster's dilemma along with potential remedies.
1. INTRODUCTION
Public forecast evaluation often concentrates on extreme events, but conditioning evaluation on extreme outcomes can discredit skillful forecasts. The paper illustrates this forecaster’s dilemma and motivates probabilistic forecasts with properly weighted scoring rules as a potential remedy.
- Extreme events attract substantial public attention, creating demand for comparative assessments of forecasts across fields including economics, finance, meteorology, and seismology.
- Restricting forecast evaluation to observed extreme cases can make always predicting calamity appear worthwhile and discredit skillful forecasts.This selection conditions evaluation on outcomes and can produce misguided judgments about predictive ability.
- In the simulation, the perfect forecast is preferred by MAE and MSE overall, but the deliberately misguided extremist forecast receives the lowest mean score among the largest 5% of observations.
- Point forecasts offer no obvious way to emphasize extreme outcomes while appropriately adapting established evaluation methods.
- Probabilistic forecasts provide an alternative because properly weighted scoring rules can emphasize extreme events while supporting comparative evaluation.
- The paper reviews forecast-evaluation theory, uses simulations, and studies probabilistic forecasts of gross domestic product growth to analyze the dilemma and possible remedies.
2. FORECAST EVALUATION AND EXTREME EVENTS
Forecast evaluation should assess the joint forecast–outcome distribution rather than condition on realized extremes. For probabilistic forecasts, proper weighted scoring rules offer a coherent way to emphasize tail regions while preserving forecast-evaluation principles.
- Joint distribution framework: Calibration concerns the conditional distribution of observations given forecasts, whereas sharpness concerns the concentration of predictive distributions.Auto-calibration implies corresponding point-forecast unbiasedness and probabilistic calibration.
- The forecaster’s dilemma: Conditioning forecast evaluation on extreme observations can make a deliberately misspecified forecast appear superior and discredit a skillful forecast.The problem arises because evaluation is stratified by the realized outcome rather than the forecast.
- Proper scoring rules: Proper scoring rules evaluate calibration and sharpness jointly by assigning penalties to predictive distributions and realized observations.They support comparative evaluation and ranking of competing probabilistic forecasts.
- Tailored scoring rules: Indicator weighting for a threshold r preserves the ranking produced by restricting evaluation to observations with y ≥ r, but the resulting weighted score is improper.The indicator rule assigns zero weight outside the target region rather than discarding observations, while producing the same comparative ranking.
- Tailored scoring rules: Proper weighted scoring rules, including threshold-weighted CRPS, emphasize extreme-event regions while retaining decision-theoretic propriety.These rules address the forecaster’s dilemma for probabilistic forecasts rather than simply conditioning on realized outcomes.
- Diebold–Mariano tests: Forecast-comparison tests use mean score differences, with variance estimation based on autocovariances of the score-difference sequence.For k-step-ahead ideal forecasts, errors are at most (k − 1)-dependent, motivating the variance estimator used for testing.
3. SIMULATION STUDIES
The simulations demonstrate that conditioning forecast evaluation on extreme outcomes can reverse rankings, especially when predictability is weak. They also assess whether proper weighted scoring rules and Diebold–Mariano tests remedy this problem, finding that benefits are not consistent at increasingly extreme thresholds.
- 3.1 The influence of the signal-to-noise ratio: The simulations extend the point-forecast experiment to probabilistic forecasts and vary σ, which controls the signal-to-noise ratio.Small σ implies high signal-to-noise and large σ implies low signal-to-noise; marginally, Y remains standard normal.
- 3.1 The influence of the signal-to-noise ratio: Under CRPS and LogS, the perfect forecast outperforms the alternatives, but restricted scores reverse the ranking and favor the misguided extremist forecast.The restricted scores use only observations exceeding 1.64, demonstrating the forecaster’s dilemma for probabilistic forecasts.
- 3.1 The influence of the signal-to-noise ratio: Proper weighted scoring rules restore the rankings obtained with unweighted CRPS and LogS when the right tail is emphasized.This provides a decision-theoretically coherent alternative to evaluating only the observed extreme cases.
- 3.1 The influence of the signal-to-noise ratio: The forecaster’s dilemma is tied to moderate or low signal-to-noise ratios, where predictability is weak.For small σ, when the signal in µ is strong, the rankings agree with those from CRPS and LogS.
- 3.2–3.3 Diebold–Mariano tests and the Neyman–Pearson lemma: In the revisited Diebold–Mariano simulations, unweighted LogS can have much higher power than weighted scores at small thresholds, although it remains below the likelihood-ratio test.The weighted twCRPS and CL tests can also exceed the nominal level α = 0.05 for some thresholds.
- 3.4 Further experiments: Across further experiments, proper weighted scoring rules do not consistently outperform unweighted scores, and their advantages vanish at increasingly extreme thresholds when the superior distribution has heavier tails.For weighted twCRPS and CSL in Scenario B, desired rejections decay to zero while undesired rejections increase with the threshold.
4. CASE STUDY
The case study compares probabilistic AR and VAR forecasts of U.S. GDP growth and inflation using real-time data, standard and threshold-weighted scoring rules. The AR-TVP-SV model generally performs best, while restricting evaluation to extreme outcomes reverses economically plausible forecast rankings.
- 4.1 Data: The study evaluates quarterly U.S. GDP growth and inflation forecasts from 1985 Q1 through 2011 Q2 using real-time data vintages.Forecasts focus on horizons of one and four quarters ahead; revised GDP observations are taken from the second available estimates.
- 4.2 Forecasting models: The forecasting models comprise constant-volatility AR and VAR schemes alongside time-varying-parameter stochastic-volatility extensions.The VAR models jointly model GDP growth, inflation, unemployment, and the three-month government bill rate.
- 4.3 Results: The AR-TVP-SV model has the best predictive performance across variables and proper CRPS and LogS scoring rules, outperforming the baseline AR model.Diebold-Mariano p-values range from 0.00 to 0.06 except for GDP growth LogS at horizon k = 4, where p = 0.37; VAR models do not outperform AR models.
- 4.3 Results: Restricting evaluation to extreme observations produces substantially different rankings and makes all models appear less skillful for current-quarter inflation than four quarters ahead.This counterintuitive reversal illustrates the forecaster’s dilemma and the danger of conditioning on outcomes.
- 4.3 Results: Proper threshold-weighted CRPS preserves rankings similar to unweighted CRPS, with AR-TVP-SV predominantly best and current-quarter forecasts more skillful than four-quarter-ahead forecasts.The associated two-sided Diebold-Mariano p-values are generally larger than under unweighted CRPS.
5. DISCUSSION
The discussion argues that conditioning forecast evaluation on realized extreme outcomes creates the forecaster’s dilemma, especially when signal-to-noise ratios are low. It recommends evaluating all cases while using proper weighted scoring rules to emphasize relevant regions, although power gains are generally limited.
- 5. DISCUSSION: Restricting forecast evaluation to extreme observations can discredit even the most skillful forecasts when the data-generating process has a low signal-to-noise ratio.The paper identifies macroeconomic and seismological predictions as settings where this issue may arise.
- 5. DISCUSSION: Conditioning on realized outcomes lacks theoretical guidance for interpreting the resulting conditional distributions, whereas conditioning on forecasts is unproblematic.From the scoring-rule perspective, indicator weighting makes proper scores improper and permits hedging.
- 5. DISCUSSION: The paper identifies evaluation over all available cases as the main remedy and proper weighted scoring rules as a way to emphasize specific regions.These weighted scores retain decision-theoretic coherence for probabilistic forecasts.
- 5. DISCUSSION: The power benefits of proper weighted scoring rules are generally limited relative to standard unweighted scoring rules and vanish with increasingly extreme thresholds.Finite-sample Diebold-Mariano behavior then depends only on forecast-distribution tail properties.
- 5. DISCUSSION: Alternative extreme-focused evaluations include induced binary tail-event probabilities and low- or high-level quantile forecasts.The paper presents these as additional functionals or summaries of predictive distributions.
APPENDIX A: IMPROPRIETY OF QUADRATIC APPROXIMATIONS OF WEIGHTED LOGARITHMIC SCORES
The appendix shows that quadratic approximations of weighted logarithmic-related scores can violate propriety. Counterexamples for conditional and censored likelihood scores establish impropriety near specified forecast parameters.
- APPENDIX A: The appendix studies quadratic approximations to the conditional likelihood and censored likelihood scores under a tail-emphasizing weight function.The construction uses a standard-normal reference distribution and w(y) = 1{y ≥ 1}.
- APPENDIX A: The conditional-likelihood quadratic approximation is strictly negative near µF = 1.314 and σF = 0.252, so CLq is not proper.The contradiction is obtained under the specified choices of weight and predictive distributions.
- APPENDIX A: The censored-likelihood quadratic approximation is strictly negative near µF = 0.540 and σF = 0.589, so CSLq is not proper.This supplies a corresponding counterexample for the censored likelihood score.
APPENDIX B: ONLINE SUPPLEMENT: MEDIA ATTENTION ON EXTREME EVENTS
The online supplement presents media coverage illustrating the focus on extreme events in public discussions of forecast quality. The listed sources were accessed August 8, 2015.
- APPENDIX B: Table 8 compiles media coverage illustrating public attention to extreme events when forecast quality is discussed.The sources span the examples selected for the paper’s supplementary media-attention analysis.
- APPENDIX B: The sources in the media-coverage table were accessed August 8, 2015.