Source-linked AI summary
Dynamic Regime-Aware Conformal Calibration for Reliable Economic Forecast Intervals under Multiple Distribution Shifts
Bogdan Oancea
TL;DR
Economic forecasting often violates conformal prediction’s exchangeability assumption across changing regimes and other distribution shifts. DRACP combines regime-aware weighting with online calibration control, achieving the closest nominal coverage across horizons and surge conditions while sacrificing interval efficiency.
Problem
Existing conformal methods face multiple concurrent regimes that a method designed for one mechanism cannot correct.
Method
DRACP integrates density-ratio, localized-kernel, probabilistic regime-aware weighting with an online significance controller in weighted conformal calibration.
Results
Coverage was 0.890 against a nominal 0.90, closest on the panel, held across forecast horizons and base forecasters, and never fell below 0.80 on any series.
Takeaways & Limitations
DRACP prioritizes reliable calibration over the sharpest intervals when forecast intervals must satisfy coverage standards.
Takeaways & Limitations
DRACP is the most expensive procedure considered, at roughly six times the cost of split conformal prediction.
Abstract
from arXiv · showhide
Conformal prediction provides distribution-free prediction intervals but relies on exchangeability, an assumption often violated in economic forecasting because of covariate shift, concept drift, local heterogeneity and latent regimes. We propose Dynamic Regime-Aware Conformal Prediction (DRACP), which combines density-ratio, localized kernel and probabilistic regime-aware weighting with a self-tuning online significance controller in a unified weighted conformal calibration framework. We distinguish three theoretical results: finite-sample validity under oracle importance weights, a coverage-gap bound for estimated weights with rates in effective sample size, and deterministic or regret guarantees for the online controller. We evaluate DRACP against six baselines on 48 real forecasting series covering euro-area and EU-27 HICP inflation, US macroeconomic and energy indicators, and daily financial series. Recent online methods (FACI, strongly-adaptive online conformal prediction and conformal PID) were verified against the authors' implementations. DRACP is not the most efficient method: strongly-adaptive online conformal prediction achieves the best interval score and intervals about 20% narrower. Instead, DRACP provides the most reliable calibration, achieving coverage closest to the nominal 0.90 (0.890), never falling below 0.80 on any series, maintaining the best coverage at all forecast horizons, and performing best during the 2021-2023 inflation surge. The strongly-adaptive method undercovers on 20 of 48 series versus 10 for DRACP. DRACP therefore offers a principled trade-off between calibration and efficiency, favoring reliable coverage when prediction intervals must satisfy coverage standards. An ablation study shows that the online controller and conditional-scale normalization provide most of the performance gain, whereas the weighting components make a smaller contribution.
1 Introduction
The introduction motivates DRACP as a unified response to simultaneous distribution shifts in economic forecasting and outlines its theoretical guarantees and broad empirical evaluation. It emphasizes a trade-off: DRACP prioritizes reliable, stable coverage over interval sharpness.
- Motivation: Economic forecasting violates exchangeability through covariate drift, structural breaks, volatility clustering, and movement between latent regimes.Real series can exhibit seasonal shift, abrupt changes, time-varying volatility, and recurring regimes simultaneously.
- Method: DRACP combines density-ratio, kernel, probabilistic regime-similarity, and online miscoverage-level weighting in one weighted-conformal calibration step.The design targets covariate shift, local heterogeneity, latent structure, and changing coverage feedback together.
- Theory: The theory establishes finite-sample coverage under oracle weights, an interpretable robustness bound, and a long-run coverage property for the online controller.The robustness bound decomposes coverage error into interpretable perturbation terms.
- Evaluation: 48 real series are used to benchmark DRACP against six baselines spanning inflation, US macroeconomic and energy indicators, and daily financial data.Recent online procedures were implemented according to published algorithms and validated against the authors’ reference code.
- Results: DRACP favors reliable coverage over sharpness, remaining stable across horizons, point forecasters, and the 2021–23 inflation surge.The evaluation presents a trade-off rather than dominance, while ablation analysis identifies where reliability comes from.
2 Related work
Related work addresses exchangeability violations through covariate-shift weighting, localization, score normalization, and online level control, but these mechanisms typically target separate departures. DRACP composes them with regime weighting and a self-tuning controller in one analyzable calibration framework.
- Split conformal prediction: Split conformal prediction uses held-out conformity-score quantiles and provides finite-sample, distribution-free, model-agnostic validity under exchangeability.Its equal weighting of calibration points makes departures from exchangeability translate into coverage error.
- Weighted and localized conformal methods: Covariate-shift conformal prediction restores validity with likelihood-ratio weighting under known density ratios, but estimated ratios degrade guarantees and cannot repair conditional-distribution changes.Localized conformal prediction instead uses query proximity, while normalized scores and quantile regression adapt through conditional scale or quantile estimates.
- Weighted and localized conformal methods: Localization trades bias for variance, so DRACP selects the kernel bandwidth adaptively subject to an explicit effective-sample-size floor.Small bandwidths adapt sharply but destabilize the weighted quantile.
- Sequential and online conformal inference: Online conformal methods control sequential miscoverage through adaptive levels, expert aggregation, subinterval regret, or proportional-integral-derivative updates.ACI uses a single step size, FACI aggregates multiple rates, SAOCP uses geometrically growing expert lifetimes, and conformal PID acts directly on interval radius.
- Economic applications and evaluation: Economic applications have used conformal intervals, but systematic comparison across a broad macroeconomic panel against modern online baselines has been missing.The paper evaluates methods with the interval Winkler score, which jointly penalizes width and miscoverage.
- DRACP’s positioning: DRACP combines covariate, localized, and regime-aware weighting with a self-tuning controller because economic series can experience simultaneous shifts whose estimation errors interact.Theorem 1 makes the resulting coverage-gap contributions analyzable within a single weighted-quantile calibration.
3 Dynamic regime-aware conformal prediction
DRACP extends split conformal prediction by weighting past errors according to recency, covariate similarity, local proximity, and latent-regime similarity, then adapting the significance level online. The procedure is causal and combines relevance weighting with controller-based correction, while fallback safeguards improve numerical stability without restoring exact validity under drift.
- Weight construction: DRACP replaces equal treatment of past errors with a normalized product of recency, density-ratio, local-kernel, and regime-similarity weights.Recency discounts older observations; density ratios target covariate shift; kernels localize predictors with an effective-sample safeguard; regime weights match latent states.
- Weighted calibration: The prediction interval is the weighted quantile of past errors, with the four relevance components multiplied and normalized before calibration.The weighted quantile replaces the unweighted conformal quantile and incorporates a test-point mass in the formal calibration step.
- Online adaptation: The online controller adjusts the working significance level from recent realized coverage, complementing weights that determine which past errors are used.The weights respond to changing conditions, whereas the controller changes which quantile is selected; the update drives long-run empirical miscoverage to α regardless of weight quality.
- Causal implementation: All quantities are computed causally using information available before each forecast origin, with nuisance quantities re-estimated from preceding data.No future observation influences any interval.
- Safeguards and limitations: Fallback to globally weighted calibration improves numerical stability when localized weights are too concentrated but does not restore exact conformal validity.The method falls back when effective sample size is below its floor or a single weight exceeds a cap; exact validity is unavailable under concept drift, estimated weights, and serial dependence.
4 Theoretical properties
The theoretical results separate ideal weighted-conformal validity from estimation-error coverage bounds and online-control guarantees. They also establish effective-sample-size protection, approximate regime-conditional coverage, and explicit limitations under dependence and fixed temporal decay.
- Online controller: The self-tuning controller has average-miscoverage and no-regret guarantees, with regret R_T=O(√T log K) against the best fixed expert and R_T/T→0.The guarantee concerns the specified per-expert pinball loss and does not by itself establish period-by-period coverage.
- Oracle validity: Oracle importance weights yield finite-sample coverage at least 1−α_t under conditional exchangeability and distinct scores.This validates the weighted-conformal construction in the idealized case, not the implemented DRACP weights.
- Estimated weights: Theorem 1 bounds implemented coverage error through additive estimation terms for each weighting component, stochastic error shrinking with effective calibration size, and irreducible conditional drift.The bound’s stochastic term depends on effective rather than nominal sample size; under dependence, effective size is reduced to roughly ESS_t/τ.
- Asymptotic scope: With fixed temporal decay λ<1, effective sample size converges to (1+λ)/(1−λ), so the coverage bound does not vanish; vanishing gaps require λ_m→1 with ESS_t→∞.This is identified as a limitation of the implemented procedure rather than an artifact of the proof.
- Regime and ESS guarantees: Approximate regime-conditional coverage degrades with posterior-estimation error and regime overlap, while deterministic ESS floors prevent the calibration size from collapsing.The ESS mechanism requires no data assumptions and guarantees a non-vacuous bound at every sample size.
5 Data and experimental design
The evaluation combines 48 real forecasting series with synthetic processes isolating major departures from exchangeability. Experiments use common point forecasts and chronological splits, compare DRACP with six baselines, and assess score, coverage, width, and uncertainty across repeated runs.
- Synthetic design: Seven synthetic data-generating processes isolate stationarity, covariate shift, conditional and abrupt or gradual drift, regime switching, and their combination.They share a nonlinear process and vary the evolution of the intercept, innovation scale, covariate distribution, and latent regime; these scenarios are diagnostic rather than primary evidence.
- Common forecasting setup: All methods use the same fitted point forecaster, chronological training, calibration, and test blocks, isolating differences in the calibration layer.The histogram-based gradient-boosting regressor is trained once and then held fixed, so sequential updating concerns calibration rather than refitting.
- Comparators and metrics: DRACP is compared with six baselines: split, rolling, ACI, FACI, SAOCP, and conformal-PID online procedures.The primary metric is mean interval (Winkler) score, supplemented by empirical coverage and mean width.
- Repetition and uncertainty: 20 repeated seeds support evaluation, while real-series uncertainty additionally uses a paired circular moving-block bootstrap with 2000 replicates and per-series Diebold–Mariano tests.Friedman/Nemenyi and Wilcoxon analyses assess consistency of method rankings across series.
6 Results
Across 48 real series, DRACP sacrifices interval-score efficiency for stronger calibration: it ranks third overall but stays closest to nominal coverage and avoids severe undercoverage. This trade-off persists across forecasters, horizons, and synthetic shifts, while weighting contributes little to score performance.
- Overall ranking: Strongly-adaptive online conformal prediction ranks first in interval score at 2.17, while DRACP ranks third at 3.15 and is beaten head-to-head on 40 of 48 series.DRACP is best on 5 series; FACI ranks second at 2.71.
- Calibration: 0.890 is DRACP’s mean coverage against nominal 0.90, closest among seven methods, with mean absolute deviation 0.027.DRACP and FACI are the only methods never below 0.80 on any series; DRACP’s worst case is 0.808.
- Forecaster robustness: DRACP’s coverage stays within 0.889–0.906 across forecasters, whereas strongly-adaptive coverage ranges from 0.863 to 0.891 and undercovers under every forecaster.Strongly-adaptive ranks first in interval score under all five forecasters, while DRACP’s score is worse than the best baseline in every case.
7 Discussion
DRACP’s main contribution is reliable coverage across data cuts, horizons, forecasters, and the 2021–23 inflation surge, while sacrificing interval sharpness. Its benefits are conditional on the forecasting setting, horizon, and point forecaster, motivating a targeted deployment rule.
- Coverage reliability: 0.890 mean empirical coverage versus nominal 0.90 is closest among seven methods, with no series below 0.80 and the highest coverage at all horizons.Its margin over strongly-adaptive online conformal prediction widens from 2.2 to 5.6 coverage points between h = 1 and h = 12, and it leads during the 2021–23 inflation surge.
- Efficiency trade-off: 2.17 versus 3.15 in rank, strongly-adaptive online conformal prediction wins head-to-head on 40 of 48 series with intervals about a fifth narrower.DRACP ranks third on interval score, so its coverage advantage comes with lower efficiency.
- Horizon dependence: By twelve steps, static calibration outperforms every adaptive method, while at horizons h ≥3 a purely online controller dominates because origin-based weights lose relevance.Horizon-specific weighting recovers only part of the multi-step loss.
- Ablation: 12% of interval score is explained by each of conditional-scale normalisation and the online controller, versus 6% combined for four weighting mechanisms.Removing the controller raises mean absolute coverage deviation from 0.027 to 0.051, whereas removing conditional scale leaves it at 0.030.
- Limitations: DRACP is beaten by simple baselines on near-efficient-market series, costs roughly six times split-conformal prediction, and its advantage disappears under the best-performing linear autoregression.The efficiency comparison depends on the point forecaster, and gains should be attributed to scale normalisation plus weighting rather than weighting alone.
- Deployment rule: DRACP is appropriate when interval coverage, rather than sharpness, is the criterion under simultaneous and unidentified shifts; outside that domain, a simpler procedure is preferable.The stated deployment rule is deliberately conditional rather than a general recommendation.
8 Conclusion
DRACP unifies density-ratio, localization, and regime weighting with a self-tuning controller and establishes validity, coverage-gap, miscoverage, and regret guarantees. Empirically, it prioritizes calibration over interval sharpness, with reliability driven mainly by the controller and open questions around width and joint weight learning.
- Method and theory: DRACP combines density-ratio, localization, and regime weighting with a self-tuning significance controller in one weighted-conformal calibration step.The framework integrates these components rather than applying them as separate calibration procedures.
- Method and theory: Finite-sample validity holds under oracle importance weights, while estimated weights receive a coverage-gap bound in effective sample size with a deterministic floor.The theoretical guarantees also include average-miscoverage control for a single-rate controller and a regret bound for the multi-rate default.
- Empirical trade-off: Third on interval score, DRACP produced intervals about a fifth wider than the sharpest competitor, illustrating its calibration–sharpness trade-off.The sharper method fell below 0.80 on five series, whereas DRACP did not.
- Empirical trade-off: 0.890 coverage against a nominal 0.90 was closest on the panel, held at every forecast horizon and under every base forecaster.Coverage remained reliable through the 2021–23 inflation surge and never fell below 0.80 on any series.
- Ablation and open problems: The ablation identifies the controller, rather than composite weighting, as the main source of reliability; open problems concern reducing width and jointly learning weights.Composite weighting was the smallest of the three ingredients by both score and coverage.
Appendix A Proofs
The appendix establishes the paper’s theoretical guarantees in sequence: oracle weighted-conformal validity, estimated-weight coverage-gap bounds, online-controller guarantees, and regime-aware and selection-hedge results. It also delineates implementation limits, including dependence, marginal-versus-conditional coverage, and the absence of aggregate deterministic coverage guarantees.
- Oracle validity: Proposition 1 proves finite-sample validity for weighted conformal prediction under oracle importance weights, but this guarantee does not automatically extend to implemented product weights.The proof uses weighted exchangeability under the true likelihood ratio.
- Estimated-weight bounds: Theorem 1 decomposes the coverage gap into weighting, regime, localization, drift, sampling, and quantile-discretization terms, with sampling controlled by effective sample size.The proof notes that dependence requires a mixing-based concentration inequality and that the stated coverage is marginal over (Xt, Yt), not conditional on Xt.
- Online control: Theorem 2 establishes bounded adaptive levels and an average-miscoverage guarantee, while exponential-weights aggregation achieves regret against the best fixed learning-rate expert.The aggregate level does not inherit the single-recursion deterministic coverage property, so part (b) is confined to pinball regret.
- Regime and selection guarantees: Propositions 3 and 4 quantify posterior-regime weighting error and provide prediction-with-expert-advice guarantees for DRACP-Select’s Hedge aggregation.For overlapping regimes, the separation-error term does not vanish with posterior estimation error alone.
Appendix B Reproducibility … C.2 The interface
The appendices specify an end-to-end reproducibility protocol, released software artefacts, installation requirements, and a scikit-learn-style interface for applying DRACP to user-supplied forecasting models. Experiments use chronological, shared protocols with repeated fits and paired bootstrap inference, while the interface supports online calibration and configurable DRACP components.
- Appendix B Reproducibility: Public series are downloaded programmatically from Eurostat, FRED, and the UCI ElectricityLoadDiagrams archive, then transformed and represented with lag, rolling-statistic, and seasonal features.Transformations include levels, first differences, or log-differences; daily series are capped at recent observations.
- Appendix B Reproducibility: Chronological train, calibration, and test blocks use no shuffling or future information, with a shared forecaster, splits, random seeds, and 20 Monte-Carlo repetitions.Synthetic seeds redraw the data-generating process, whereas real-series seeds reseed stochastic estimators.
- Appendix B Reproducibility: 96 of 288 paired comparisons significantly favor DRACP on mean interval score, 21 favor it against DRACP, and 171 show no separation at 5%.The 288 comparisons cover six baselines across 48 series; DRACP is significantly better 12–19 times per baseline and worse 2–5 times.
- Appendix B Reproducibility: The released MIT-licensed Python package dracp includes an archived version, installation support, and a run.sh entry point that reproduces data preparation, experiments, tests, tables, and figures.Raw predictions, per-seed metrics, bootstrap intervals, and aggregate summaries are saved for recomputation without rerunning experiments.
- Appendix C Using the software for your own forecasts: DRACP can calibrate any trusted point-prediction model because conformal validity does not depend on the forecasting model.The appendix provides installation steps and a complete worked example for user-supplied forecasts.
- C.1 Installation: Python 3.10 or later and NumPy, pandas, SciPy, and scikit-learn are required, with installation available through the Python Package Index or editable source checkout.A one-line import-and-print command checks whether installation succeeded.
- C.2 The interface: ConformalForecaster accepts any model with .fit and .predict, fits chronologically, calibrates on recent data, reveals outcomes sequentially, and reports coverage, width, and interval score.Configuration fields such as calibration_window, n_regimes, bandwidth, use_faci_control, and asymmetric can be passed as keywords.
C.3 A complete example … Appendix D Data-generating processes and series descriptions
The appendix provides a reproducible DRACP example, variant configurations, practical guidance, and documentation of synthetic and real evaluation series. The example attains 0.94 empirical coverage against nominal 0.90, while guidance highlights adaptation trade-offs and limits on near-random-walk series.
- C.3 A complete example: 0.94 empirical coverage is achieved against nominal 0.90, with intervals widening around the structural break in the example.The example uses a monthly series with a level shift and volatility change at t = 300.
- C.3 A complete example: Passing y to predict_interval enables online adaptation, whereas omitting y yields fixed calibration.Sequential methods consume outcomes and mutate calibration state, so the test block should be evaluated in one pass.
- C.4 Accessing the variants: The default configuration combines composite weighting, the self-tuning controller, and symmetric scores.The self-tuning controller is explicitly enabled in the paper’s experiment configurations.
- C.4 Accessing the variants: DRACP-Select runs several predictors in parallel and can use hard selection or convex averaging of constituent radii.The averaging mode is selected with mode="average".
- C.5 Practical guidance: Shorter calibration windows adapt faster but reduce effective sample size, while two or three regimes typically represent economically distinct states.The method automatically enforces an effective-sample-size floor and exposes estimated regime posteriors for inspection.
- C.5 Practical guidance: On near-random-walk series such as daily asset returns, simpler rolling calibration performs at least as well at a fraction of the cost.The method is designed for series subject to several simultaneous shifts.
- Appendix D Data-generating processes and series descriptions: Appendix D documents the synthetic data-generating scenarios and real series used so both evaluation components can be reproduced or challenged.The documentation covers precisely how scenarios are generated and which real series are included.
D.1 Synthetic data-generating processes
The seven synthetic scenarios share a nonlinear regression backbone but vary in intercept, innovation scale, covariate distribution, and latent-regime evolution. With n = 3000 and a 45%/25% training-calibration split, the designs span stationary data, distribution shifts, structural drift, regime switching, and their combination.
- Common backbone: All scenarios use five-dimensional standard-normal covariates before perturbation and a nonlinear response with sine and squared-covariate terms plus Gaussian noise.The nonlinear conditional mean deliberately misspecifies a linear forecaster and creates a non-trivial residual distribution for conformal calibration.
- Experimental protocol: 45% of n = 3000 observations are used for training, 25% for calibration, and the remainder for sequential evaluation.Normalised time is defined as ϕ_t = t/(n −1).
- Stationary control: The stationary control sets c_t = 0 and σ_t = 1 without perturbation, so exchangeability and the classical guarantee apply.This scenario is explicitly designated as the control case.
- Drift scenarios: Abrupt and gradual drift vary level, covariate scale, and volatility either through a break at t = 0.65n or through continuous time evolution.Abrupt drift changes c_t from 0 to 4 and σ_t from 1 to 1.8; gradual drift uses c_t = 2 sin(2πϕ_t), covariate scaling 0.5 + 1.5ϕ_t, and σ_t = 0.7 + ϕ_t.
- Regime switching: Regime switching uses a persistent three-state Markov chain that jointly determines intercept, innovation scale, and covariate shifts.The regimes recur and represent expansion- and recession-like states.
D.2 Real forecasting series
The evaluation uses 48 real forecasting series across HICP inflation, US macroeconomic and energy indicators, and daily financial and interest-rate changes. Series are modelled one step ahead with chronological, leakage-free splits, while revisions and rolling daily cutoffs limit exact reproduction.
- Modelling and evaluation: Each series is modelled one step ahead using autoregressive lags, rolling statistics, and sine/cosine seasonal encodings at its natural period.Seasonal periods are 12 for monthly series and 7 for daily series; transformations vary by series type.
- Modelling and evaluation: Chronological splits use no shuffling and no future information at any point.Growth series use log-differences, rates and spreads use first differences, and unemployment and inflation rates remain untransformed.
- Reproducibility caveats: Monthly macroeconomic series are subject to statistical revision, so later downloads may differ slightly for recent observations.Daily observations are capped to keep sequential evaluation tractable, causing the evaluation window to advance with the download date and affecting only the daily group.
- Dataset composition: 48 series span three blocks: 28 euro-area and EU-27 HICP inflation series, US macroeconomic and energy indicators, and seven daily financial and interest-rate series.The HICP panel covers the 2021–23 inflation surge; the financial block is a deliberate negative control with little exploitable shift structure.
Appendix E Full numerical results
Appendix E provides complete per-series results for all 48 real series, linking interval score, coverage, width, and Diebold–Mariano comparisons to the aggregate analyses. The tables show that DRACP’s score advantage reflects calibration and avoidance of miss penalties rather than narrower intervals.
- Real-series results: DRACP remains close to the nominal 0.90 coverage throughout, so low score alone is not treated as sufficient evidence of competitiveness.The appendix explicitly pairs empirical coverage with interval score because narrow but poorly calibrated intervals are not valid competitors.
- Real-series results: DRACP’s interval widths are typically comparable to or slightly larger than online baselines and far larger than split conformal prediction.Its advantage is attributed to avoiding the large miss penalties incurred by narrow but miscalibrated intervals.
- Real-series results: Significant Diebold–Mariano results concentrate on structured series, while adverse cases cluster in the financial group.Table 20 contains one DRACP-versus-baseline test for each series–baseline pair, with significance assessed using unadjusted and Benjamini–Hochberg-controlled thresholds.
- Provenance and synthetic results: Per-series score, coverage, and width results average 20 repeated fits, whereas the Diebold–Mariano matrix uses a single fit with seed 42.The synthetic table averages 20 Monte Carlo repetitions with a redrawn data-generating process; series are grouped into EU-27 HICP, US macroeconomic and energy, and daily financial blocks.
Appendix F Additional figures
Appendix F provides additional figures covering per-series performance, internal DRACP diagnostics, robustness, computational cost, and prediction-interval examples. The diagnostics include quantities identified by Theorem 1 as controlling the coverage gap and therefore serve as deployment diagnostics.
- Internal diagnostics: Figures 14 and 15 show density-ratio contributions, posterior regime probabilities, the online significance level α_t, effective calibration sample size, transition matrices, localization weights, and residual diagnostics.These quantities are identified by Theorem 1 as controlling the coverage gap and thus function as deployment diagnostics.
- Robustness and cost: Figure 16 examines sensitivity to localization bandwidth, temporal forgetting factor, and density-ratio clipping in a combined-shift scenario.These are the robustness hyperparameters evaluated for DRACP.
- Robustness and cost: Figures 17 and 18 cover coverage under increasing shift complexity, interval-width distributions under combined shift, runtime versus calibration-buffer size, and O(Bd) memory scaling.The computational figure reports runtime per prediction and the memory footprint separately.
- Example intervals: Figures 19 and 20 provide example DRACP prediction intervals for a synthetic regime-switching series, electricity demand, and euro-area inflation.The synthetic example annotates regimes and uses the combined-shift scenario.
- Per-series performance: Figures 12 and 13 report per-series mean interval score and empirical coverage, with repeated-fit or seed means and 95% bootstrap uncertainty, for representative structured and financial series.Figure 13 marks nominal 0.90 coverage with a dashed reference line.