Source-linked AI summary

Autoregressive Process Modeling via the Lasso Procedure

Yuval Nardi, Alessandro Rinaldo

arXiv:0805.1179v1math.ST

TL;DR

The paper addresses how to fit autoregressive time-series models when the maximal lag grows with sample size, rather than assuming a fixed low-dimensional order. It formulates a Lasso estimator in this double asymptotic setting and derives consistency results under stated conditions. For p = O(log n), the estimator achieves model-selection, estimation, and prediction consistency.

  • Problem

    Classical time-series modeling commonly assumes fixed, low-dimensional orders, while this paper studies autoregressive modeling when the maximal lag grows with sample size.

  • Method

    The paper applies a penalized ℓ1 Lasso estimator to autoregressive coefficients under a double asymptotic framework with p = O(log n).

  • Results

    For p = O(log n), the Lasso estimator is model selection consistent, estimation consistent, and prediction consistent.

  • Takeaways & Limitations

    The fitted model can be selected among AR models with maximal lag from 1 through O(log n), approximating a general linear time series as n grows.

  • Takeaways & Limitations

    The theoretical investigation assumes Gaussian innovations, although the paper states that this assumption can be relaxed.

Abstract

from arXiv · show

The Lasso is a popular model selection and estimation procedure for linear models that enjoys nice theoretical properties. In this paper, we study the Lasso estimator for fitting autoregressive time series models. We adopt a double asymptotic framework where the maximal lag may increase with the sample size. We derive theoretical results establishing various types of consistency. In particular, we derive conditions under which the Lasso estimator for the autoregressive coefficients is model selection consistent, estimation consistent and prediction consistent. Simulation study results are reported.

1 Introduction

The paper applies the Lasso to autoregressive time-series modeling when the maximal lag grows with sample size. It develops asymptotic results and supports the approach with simulations.

  • Classical ARMA modeling often relies on known, fixed, low-dimensional orders that are difficult to verify in practice.
  • The Lasso simultaneously performs model selection and estimation for linear models.
  • The paper studies a double asymptotic framework in which the number of autoregressive parameters grows with sample size.
  • The increasing-order autoregressive process lies between fixed-order and infinite-order AR processes and includes many ARMA processes.
  • The paper derives model-selection, estimation, and prediction consistency results and reports a simulation study.

2 Penalized autoregressive modeling

The paper defines a penalized ℓ1 estimator for autoregressive coefficients in a growing-lag framework. It uses p = O(log n), allowing simultaneous selection and estimation across candidate AR models.

  • The model assumes observations follow an AR(p) process driven by independent Gaussian innovations.
  • The process is assumed causal, with an absolutely convergent MA(∞) representation and roots outside the unit-disc boundary condition.
  • The Lasso estimator minimizes penalized ℓ1 least squares for the autoregressive coefficients using a design matrix of lagged observations.
  • The tuning scheme combines a grand parameter with lag-specific parameters, allowing different penalties across predictors and producing sparse solutions.
  • Choosing p = O(log n) yields favorable asymptotic properties while allowing the selected model to range over relevant AR processes.

3 Asymptotic Properties of the Lasso

The paper establishes asymptotic properties of the Lasso estimator for autoregressive models whose lag dimension grows as p = O(log n), including model-selection, estimation, and prediction consistency.

  • Model Selection Consistency: Model selection consistency concerns recovering the sparsity structure of the true autoregressive parameter, with sign consistency requiring matching coefficient signs with probability tending to one.Matching supports is a weaker form of model selection consistency implied by sign consistency.
  • Model Selection Consistency: Under Theorem 3.1’s assumptions and p = O(log n), the Lasso estimator is sign consistent.The assumptions include bounded inverse autocovariance structure and an incoherence condition controlling correlations between relevant and irrelevant variables.
  • Estimation Consistency: Estimation consistency means ||φ̂_n − φ*|| converges to zero, and under the stated conditions the estimator achieves rate O(α_n).The theorem uses p = O(log n) and a penalty condition involving λ_n and the relevant-coordinate penalties.
  • Prediction Consistency: The prediction analysis uses stationary Gaussian processes in Hρ(l, L), whose strong-mixing coefficients, AR coefficients, and autocovariances decay exponentially.The class is used to exploit additional time-series structure in establishing prediction consistency.
  • Prediction Consistency: Prediction consistency concerns convergence of fitted future-value predictions, and Theorem 3.3 provides bounds that yield a vanishing error probability under suitable penalty and sparsity conditions.Corollary 3.4 takes λ_n = n^-α with α ∈ (2/5, 1/2), while requiring (s/p)^1/2 ≤ Dn^(α−2/5).

4 Illustrative Simulations

Simulations assess Lasso variable selection for a sparse autoregressive process and compare its selected model with conventional autoregressive fitting. Across 1000 replications, cross-validation typically selected a compact set and prioritized the largest coefficients.

  • Simulation design: The simulated length-1000 process has nonzero autoregressive coefficients at lags 1, 3, 5, 10, and 15, with Gaussian innovations of standard deviation 0.1.The coefficients are 0.2, 0.1, 0.2, 0.3, and 0.1, respectively.
  • Lasso selection: For p = 50, cross-validation selected the Lasso penalty, and variables whose solution paths reached that threshold were declared significant.The simulation used one common penalty for all autoregressive coefficients.
  • Lasso selection: In the illustrative series, the Lasso included all nonzero autoregressive coefficients, while their entry order closely followed coefficient magnitude.φ10 and φ5 entered almost immediately; φ3 and φ15 entered last.
  • Comparison with conventional fitting: The conventional Yule-Walker fit was non-sparse, whereas AIC correctly estimated the autoregressive order as 15 in the illustrative comparison.The fitted values shown for the first 30 coefficients were compared against the true nonzero lags.
  • Repeated simulations: Across 1000 simulations, the number of selected variables had mean 6.42 and standard deviation 2.44, with minimum 3, median 6, and maximum 22.φ10 and φ5 were always included, whereas φ3 and φ15 had smaller but still high selection frequencies.
  • Repeated simulations: The solution paths usually selected φ10 and φ5 first and second, while φ15 and φ3 entered later and fell outside the first five selections in 1.9% and 20.2% of cases.The path-entry analysis treats earlier entry as indicating greater significance among the nonzero variables.

5 Discussion

The paper concludes that the Lasso remains consistent when the maximal autoregressive lag grows as p = O(log n), while identifying nonlinearity and extending the method to nonlinear autoregression remain important boundaries.

  • Conclusions: When p = O(log n), the Lasso estimator is model selection, estimation, and prediction consistent.The fitted model is selected among AR models with maximal lag between 1 and O(log n).
  • Conclusions: The growing-lag Lasso can virtually approximate a general linear time series as n tends to infinity.
  • Assumptions: The model-selection consistency proof does not require Gaussian noise, although Gaussianity is assumed elsewhere for theoretical development.
  • Future directions: Nonlinear autoregressive extensions are challenging because even mild p can produce very many interaction terms.The paper identifies understanding nonlinear autoregressive processes as a prerequisite for applying the Lasso reliably.

6 Proofs

The proofs establish the Lasso’s consistency results through convex optimality conditions, concentration of the autoregressive Gram matrix, and bounds on stochastic and penalty terms.

  • Lasso optimization: The Lasso estimator is defined as an optimal solution to a convex minimization problem involving the objective M_Λn(φ).The least-squares gradient and Hessian depend on the sample Gram matrix.
  • Model-selection proof: Optimality conditions characterize the estimator through subgradients, relevant and non-relevant variables, and the block decomposition of the design matrix.
  • Model-selection proof: The model-selection proof shows that the key events A and B occur with probability tending to one under the theorem’s conditions.These events encode the elementwise optimality requirements used to establish sign consistency.
  • Probability bounds: The proofs control stochastic terms using martingale moment inequalities, independence, Cauchy–Schwarz, Burkholder’s inequality, and Markov’s inequality.
  • Estimation consistency: The estimation proof concludes that ∥φ̂_n − φ*∥ = O_P(α_n) after showing the objective increases uniformly outside a sufficiently large neighborhood.
  • Estimation consistency: The normalized Gram matrix converges to the autocovariance matrix even when p grows, supporting the growing-dimensional consistency analysis.The argument bounds covariance terms and obtains an O(p^2/n) remainder.
Loading 0805.1179v1…