Source-linked AI summary
Achieving Shrinkage in a Time-Varying Parameter Model Framework
Angela Bitto, Sylvia Frühwirth-Schnatter
TL;DR
TVP models can overfit when coefficients that are actually static are allowed to vary. The paper applies double gamma shrinkage to process variances and develops an ASIS-enhanced MCMC framework for univariate and multivariate TVP models. Its applications suggest that the approach avoids overfitting when coefficients are static or insignificant.
Problem
TVP models risk overfitting because coefficients that are constant or insignificant may be modeled as time-varying, reducing statistical efficiency.
Method
The paper places double gamma shrinkage priors on process variances and combines centered and non-centered parameterizations through ASIS-based MCMC.
Results
The applications suggest that double gamma priors successfully avoid overfitting when coefficients are static or insignificant, in univariate and multivariate TVP models.
Takeaways & Limitations
The framework provides a general approach for introducing sparsity into TVP and other state-space models.
Takeaways & Limitations
Cholesky-based inference is not invariant to reordering indices, and functionals of Σ−1 showed sensitivity to data ordering.
Abstract
from arXiv · showhide
Shrinkage for time-varying parameter (TVP) models is investigated within a Bayesian framework, with the aim to automatically reduce time-varying parameters to static ones, if the model is overfitting. This is achieved through placing the double gamma shrinkage prior on the process variances. An efficient Markov chain Monte Carlo scheme is developed, exploiting boosting based on the ancillarity-sufficiency interweaving strategy. The method is applicable both to TVP models for univariate as well as multivariate time series. Applications include a TVP generalized Phillips curve for EU area inflation modelling and a multivariate TVP Cholesky stochastic volatility model for joint modelling of the returns from the DAX-30 index.
1 Introduction
TVP models flexibly capture gradual changes but can overfit when many coefficients are actually constant. The paper introduces shrinkage for process variances to distinguish time-varying, static significant, and insignificant coefficients.
- TVP models capture gradual changes over time and provide an alternative to models with multiple change points.
- 406 potentially time-varying coefficients in the DAX-30 application illustrate how TVP models can overfit when only a small fraction changes.
- Allowing static coefficients to vary over time causes considerable statistical-efficiency loss relative to specifying them as static a priori.
- The proposed shrinkage framework distinguishes time-varying coefficients, significant but static coefficients, and insignificant coefficients.
2 Sparse time-varying parameter models
The paper formulates TVP models as state-space regressions whose random-walk process variances govern coefficient dynamics. It replaces conventional variance priors with double gamma shrinkage that can concentrate static coefficients at zero while retaining flexibility for changing coefficients.
- A TVP model is a state-space regression in which latent regression coefficients follow independent random walks.
- Each process variance θj governs the dynamics of coefficient βjt and can shrink the coefficient path toward its fixed regression value.
- The observation equation may use homoscedastic or stochastic-volatility errors, allowing the error variance to be constant or time-dependent.
- The non-centered parametrization places unknown fixed coefficients and square roots of process variances in the observation equation while leaving state equations parameter-free.
- The double gamma prior is a normal-gamma-derived prior for process variances that generalizes the Bayesian Lasso and can produce a pronounced spike at zero.
- κ2 controls global shrinkage, while aξ controls local adaptation and smaller aξ yields thicker tails and more local adaptation.
3 MCMC Estimation
The paper develops an MCMC sampler for the hierarchical shrinkage model by combining centered and non-centered parameterizations through ASIS. This interweaving improves posterior sampling efficiency while accommodating process-variance shrinkage.
- The sampler uses latent-state data augmentation for both centered and non-centered TVP parameterizations, with additional latent log volatilities for stochastic volatility models.
- ASIS interweaves centered and non-centered parameterizations, combining their advantages and increasing posterior sampling efficiency over either conventional Gibbs scheme.
- The algorithm alternates state sampling, Gaussian parameter updates, interweaving, hyperparameter updates, and prior-variance sampling.
- The interweaving step resamples fixed coefficients and process-variance square roots in the centered parameterization while preserving identical posterior distributions.
- Sampling square roots of process variances avoids boundary-space problems for small variances and improves sampler mixing.
- Process variances are sampled from generalized inverse Gaussian posteriors under the double gamma specification.
4 Comparing shrinkage priors through log predictive density scores
The paper compares shrinkage priors using log predictive density scores based on sequential one-step-ahead predictive densities. These scores aggregate observation-level predictive performance over the evaluation sample.
- Log predictive density scores evaluate and compare alternative shrinkage priors in the paper.
- The first t0 observations form a training sample, while predictive evaluation uses the remaining observations.
- Individual LPDS values are log one-step-ahead predictive densities for each observation, whereas LPDS aggregates performance over the entire time series.
- Predictive densities are approximated from posterior draws using Gaussian sum or conditionally optimal Kalman mixture approximations.
- For each posterior draw, the conditionally Gaussian state-space model supplies the exact one-step predictive density through Kalman-filter prediction.
5 Extension to multivariate time series
The multivariate extension applies shrinkage priors to potentially time-varying coefficients and process variances, identifying structural zeros, constant coefficients, and insignificant effects. Cholesky decomposition converts the multivariate stochastic-volatility system into independent TVP equations suitable for Bayesian inference.
- Sparse TVP models for multivariate time series: The multivariate TVP model allows coefficients to be structural zeros, constant values, or independently evolving random walks.Each potentially unconstrained coefficient has its own process variance, which governs whether it changes over time.
- Sparse TVP models for multivariate time series: Shrinkage priors on βij and θij identify whether coefficients are time-varying, constant, or insignificant.A zero process variance corresponds to a constant coefficient, while a constant coefficient may itself be insignificant when βij = 0.
- Sparse TVP models for multivariate time series: When observation errors are uncorrelated, row-specific priors yield independent univariate TVP models that can be estimated separately or in parallel.With correlated errors, a Cholesky decomposition restores this representation.
- The sparse TVP Cholesky SV model: The time-varying Cholesky decomposition represents the multivariate distribution as independent TVP equations with lagged components as regressors.The coefficients βij,t = −Φij,t follow random walks, while the diagonal matrix Dt captures time-varying volatilities.
- The sparse TVP Cholesky SV model: The Cholesky stochastic-volatility model combines time-varying regression coefficients with row-specific stochastic-volatility processes.Each log volatility hit follows an individual SV model with parameters μi, φi, and σ2.
6 Illustrative Application to Simulated Data
The simulation evaluates whether shrinkage priors distinguish strongly time-varying, constant but significant, and insignificant coefficients. Across simulated series, the double gamma prior provides heavier shrinkage and lower average mean squared error for non-time-varying coefficients.
- Simulation design: 100 simulated series of length T = 200 contain one strongly time-varying, one constant significant, and one insignificant coefficient.The design sets θ1 = 0.02, θ2 = θ3 = 0, β2 = −0.3, and β3 = 0.
- Coefficient classification: Posterior distributions classify coefficients as time-varying, static but significant, or insignificant through βj and θj.A bimodal posterior for θj suggests a nonzero process variance, whereas a unimodal posterior suggests θj is likely zero.
- Simulation results: Shrinkage priors detect the time-varying coefficient β1t, the constant significant coefficient β2t, and the insignificant coefficient β3t.The posterior for θj shrinks toward zero when θj = 0, and β3 is also shrunk toward zero when β3 = 0.
- Simulation results: Heavier shrinkage from the double gamma prior reduces avMSE relative to the hierarchical Bayesian Lasso, especially for the two non-time-varying coefficients.The comparison uses avMSE, avVAR, and avBIAS2 over the 100 simulated time series.
- Simulation results: Posterior paths are evaluated using pointwise quantiles under the centered parametrization against the true paths.Figure 4 compares the Bayesian Lasso and double gamma priors for one simulated time series.
7 Applications in Economics and Finance
The applications show that hierarchical double-gamma shrinkage identifies static, insignificant, and time-varying coefficients in EU inflation and DAX return models. It also improves MCMC efficiency and predictive performance relative to Bayesian Lasso and inverted-gamma alternatives.
- EU area inflation: The EU inflation application estimates 37 potentially time-varying coefficients using monthly data from February 1994 to November 2010.The specification includes inflation lags, economic predictors, and monthly dummy variables.
- EU area inflation: For 34 regression coefficients, posterior medians are shrunken toward zero, and their 95%-confidence regions indicate no significance.These coefficients include lagged inflation values and monthly dummy variables.
- EU area inflation: The four highlighted predictors are the 1-month interest rate, 1-year interest rate, M3, and unemployment rate.Their posterior densities and paths are compared under hierarchical double-gamma and Bayesian Lasso priors.
- EU area inflation: The posterior paths classify M3 and unemployment as time-varying, the 1-month interest rate as significantly nonzero but constant, and the 1-year interest rate as insignificant.The 1-month interest-rate process-variance posterior peaks at zero, while the 1-year rate is shrunken toward zero.
- EU area inflation: Interweaving substantially improves MCMC mixing and reduces inefficiency factors for the EU inflation model.The comparison uses sample paths with and without interweaving and selected inefficiency factors from Table 3.
- DAX returns: In the DAX Cholesky stochastic-volatility model, double-gamma shrinkage identifies only βi1,t, βi2,t, and βi7,t as time-varying among the illustrated coefficients.It yields sharper concentration at zero for static coefficients, smoother paths, and substantially greater efficiency than the inverted-gamma prior.
- Predictive evaluation: Predictive comparisons favor shrinkage priors over the inverted-gamma prior, while the hierarchical double-gamma prior is clearly preferable to hierarchical Bayesian Lasso for the EU data.The DAX comparison evaluates cumulative LPDS over the last 500 returns; the EU comparison uses LPDS over the last 100 time points.
8 Conclusion
The paper develops Bayesian shrinkage for TVP models, using double gamma priors on process variances to reduce overfitted time variation and an efficient ASIS-based MCMC scheme. Applications cover EU-area inflation and DAX returns, with findings suggesting that double gamma priors can avoid overfitting when coefficients are static or insignificant.
- The framework uses process-variance shrinkage to automatically reduce time-varying parameters to static ones when TVP models overfit.
- The DAX application compares posterior paths under conditionally conjugate and hierarchical double gamma priors, including heat plots of a 29×28 Cholesky factor matrix.
- The method combines normal-gamma-based double gamma priors with an efficient MCMC scheme using ancillarity-sufficiency interweaving.
- Log predictive density scores are used to compare shrinkage-prior settings and inverted gamma alternatives.
- Applications include EU-area inflation modeling and a sparse TVP Cholesky stochastic-volatility model for DAX-30 returns.
- The authors conclude that double gamma priors are successful in avoiding overfitting when coefficients are static or insignificant, and that the framework is general.
Model Framework
This section contains author and institutional information for Angela Bitto and Sylvia Frühwirth-Schnatter at WU Vienna University of Economics and Business.
- The paper is authored by Angela Bitto and Sylvia Frühwirth-Schnatter.
- The authors are affiliated with the Institute for Statistics and Mathematics and the Department of Finance, Accounting and Statistics at WU Vienna.
- The listed institution is located in Vienna, Austria.
A.1.1.1 Step (a): Sampling the latent states
Step (a) samples latent states in the non-centered TVP state-space model using Gaussian state-posterior calculations, with FFBS or the faster AWOL alternative. The banded precision structure enables efficient sampling without explicitly computing matrix inverses.
- Step (a) samples the latent state sequence conditional on known parameters using FFBS or the faster AWOL algorithm.AWOL stands for “all without a loop.”
- The latent-state posterior is multivariate normal and can be represented using a tri-diagonal precision matrix.
- The precision matrix consists of d×d submatrices, with off-diagonal blocks equal to −I_d.
- The band structure makes the Cholesky decomposition computationally inexpensive and permits back-band substitution instead of explicitly calculating L^-1.
- The subsequent expanded regression conditions on the latent states and jointly represents constant coefficients and square roots of process variances.
- The conditional posterior of the regression coefficient vector is multivariate normal under the conjugate prior.
A.1.1.3 Step (c): Sampling the prior variances
This section describes conditional sampling and predictive-density evaluation within the TVP framework, including variance updates, Kalman filtering, and comparisons of predictive-density approximations. It also documents the ECB and DAX applications used for evaluation.
- The hierarchical normal-gamma priors yield generalized inverse Gaussian conditional posteriors for process variances.
- The generalized inverse Gaussian distribution is a three-parameter distribution supported on the positive real line.
- When the average variance ξ^2 is small, the posterior expectation of κ^2 is proportional to 1/d_2, so smaller d_2 encourages stronger shrinkage of θ_j toward zero.
- The Kalman filter derives the predictive density conditional on model parameters and observed past data through prediction and correction steps.
- The conditionally optimal Kalman mixture integrates over the entire latent state process and outperforms alternative approximations in the reported comparisons.
- The naive Gaussian mixture can be very imprecise for high signal-to-noise state-space models, because its narrow components can produce many zero density evaluations and bias LPDS estimates.
- The ECB data are monthly observations from February 1994 through November 2010, while the DAX data contain roughly 2500 daily returns from September 2001 through August 2011.