Source-linked AI summary
Prediction Inference of Time Series with Standard ReLU Deep Neural Networks
Kejin Wu
TL;DR
Prediction inference for nonlinear dependent time series remains challenging, especially when fitting NLAR models. This paper develops pertinent prediction intervals using standard ReLU DNN estimators and finds consistent point predictions with asymptotically valid intervals under standard assumptions.
Problem
Prediction inference for nonlinear dependent time series remains challenging, particularly because NLAR models can be difficult to fit.
Method
The paper uses a standard ReLU DNN estimator within an additive model to construct pertinent prediction intervals.
Results
Point predictions are consistent and the corresponding prediction intervals are asymptotically valid under standard assumptions.
Takeaways & Limitations
Compared with kernel estimation, DNN-based prediction intervals are more robust to the curse of dimensionality and achieve better coverage.
Takeaways & Limitations
Order selection for prediction with kernel and DNN estimators is not trivial, motivating the need for a useful order-selection criterion.
Abstract
from arXiv · showhide
We propose a methodology based on the standard ReLU Deep Neural Networks (DNN) to make predictions and quantify their uncertainty. Classically, people rely on linear, non-linear, or non-parametric kernel methods to fit and then predict the time series. As the universal approximation ability was revealed for DNN, its application has become more and more popular for prediction tasks in various scientific areas. However, the corresponding uncertainty quantification has not been studied thoroughly. Particularly, the uncertainty in prediction will consist of two parts: (1) the future variability; (2) the estimation variability within training data. To capture both variabilities, we build the so-called pertinent prediction interval (PPI) with the DNN model estimator. We first explore the consistency property of the DNN estimator with beta-mixing dependent data. Subsequently, we show that the implied forward bootstrap series is still beta-mixing and possesses the same stationary distribution as the original time series in probability, which is a key condition to enable the PPI. Lastly, the desired PPI is built after imposing minimal conditions on the limiting distribution of predictive roots. Simulations and real-data analysis are deployed to challenge our approach with standard non-parametric methods.
1 Introduction
The paper develops statistical prediction inference for nonlinear time series using fully connected ReLU DNNs, addressing limitations of traditional models and black-box neural-network forecasting. Its forward-bootstrap procedure supports consistent multi-step point prediction and prediction intervals that capture training-estimation variability.
- Motivation and limitations: The study targets nonlinear time-series prediction because ARMA models handle only linear series, while NLAR estimation and multi-step prediction inference remain difficult.Iterative nonlinear forecasting can produce biased optimal L2 multi-step predictions, and existing numerical approaches do not deliver optimal L1 point predictions when second moments may be infinite.
- Related prediction inference: Forward-bootstrap prediction methods motivate pertinent prediction intervals that incorporate model-fitting variability and can achieve the same coverage with shorter data than classical intervals.Prior work reports less finite-sample undercoverage and finite-sample effectiveness for PPI, although it is asymptotically equivalent to classical prediction intervals.
- Gap in DNN forecasting: Existing DNN forecasting studies often treat networks as black boxes and omit statistical accuracy quantification or estimation variability from training.Although DNNs offer flexible estimation and universal approximation, complete estimation inference for DNN estimators remains under development.
- Contributions: The proposed forward-bootstrap procedure makes multi-step DNN point predictions consistent with oracle conditional means and medians, while enabling asymptotically valid naive quantile intervals and DNN-based PPI.The PPI is constructed under minimal assumptions and captures estimation variability from training the DNN.
- Method scope: The analysis focuses on fully connected ReLU DNNs, specifically the fully connected ReLU multilayer perceptron, for NLAR time series.The approach is intended to combine flexible, powerful DNN estimators with statistical prediction inference.
2 Standard ReLU Deep Neural Networks
This section defines the fully connected ReLU DNN architecture and the function class used throughout the paper. It also explains how its estimation theory relates to general DNNs and specifies the paper’s under-parametrized setting.
- Architecture: The model is a fully connected multilayer perceptron whose depth and layer-wise width determine its hidden-layer structure, with every neuron connected to the preceding layer.The input and output are treated as the zeroth and (L + 1)-th layers, respectively.
- Architecture: Each hidden layer applies the ReLU activation, and the resulting network repeatedly composes affine transformations and activations to map inputs to outputs.The paper uses ReLU partly because of its robustness to gradient vanishing.
- DNN class: The considered DNN class allows depth, equal width, and output bounds to grow with sample size, while restricting depth to L_n, width to H_n, and the sup-norm to 2B_n.The growth rates of L_n, H_n, and B_n are discussed later in the paper.
- Relation to general DNNs: Because any sparse general DNN can be embedded in a fully connected MLP by zeroing missing connections, fully connected error bounds also control embedded general-DNN estimators.The fully connected class is larger and has a larger estimation-error bound, while its asymptotic results carry over to general DNN estimators.
- Parameterization and scope: The paper focuses on under-parametrized situations and reports that DNN estimators are more robust to the curse of dimensionality than standard kernel estimators.Formal inference for DNN-type estimators is described as usually requiring n > Pdim(N_n), whereas this work focuses on the under-parametrized case.
3 Point Predictions and Prediction Intervals with DNN
This section develops DNN point predictions, naive quantile prediction intervals (QPI), and pertinent prediction intervals (PPI) for general autoregressive time series. Under smoothness and dependence conditions, it establishes DNN estimation and prediction consistency, QPI validity, and PPI validity with improved finite-sample coverage.
- DNN estimation: Theorem 3.1 establishes that standard fully connected ReLU DNNs achieve vanishing optimal sup-norm approximation error, governed by sample size and target-function smoothness.The result supports DNN sieve estimation on an expanding domain, provided the domain bound B_n grows sufficiently slowly.
- Point predictions and QPI: Under Theorem 3.3 conditions, Algorithm 1 yields consistent L2 and L1 point predictions, consistent quantile predictions, and an asymptotically valid QPI.The conditional CDF estimator is uniformly consistent, and these conclusions hold for multi-step prediction under the stated domain restrictions.
- PPI construction: Although asymptotically valid, QPI can under-cover; PPI addresses this by incorporating estimation variability alongside future innovation variability.Its predictive-root construction can attain QPI-level coverage with a smaller sample and reduce finite-sample undercoverage.
- Bootstrap validity: The forward bootstrap series preserves beta-mixing and converges to the original stationary distribution, yielding PPI feasibility with the DNN estimator under minimal additional assumptions.Theorem 3.8 establishes that the bootstrap predictive root has the same non-trivial limiting distribution as the original predictive root.
4 Simulations
Simulations compare DNN- and kernel-based QPI/PPI methods across models, dimensions, sample sizes, and five-step prediction horizons. DNN point predictions are more accurate, while DNN-based PPI generally offers stronger coverage, particularly in challenging settings.
- Simulation results: DNN point predictions are more accurate than kernel predictions, especially for Models 2 and 3 when dimension d is relatively large.The simulations use sample sizes n = 50, 200, 500 and evaluate point prediction with MSE.
- Simulation results: PPI intervals are generally longer and achieve higher coverage rates than QPI intervals for both DNN and kernel estimators.PPI also provides better coverage than QPI, especially when the sample size is small.
- Simulation results: For Model 3 with n = 500, one-step PPI coverage is 89.8% with the kernel estimator versus 95.5% with the DNN estimator.Kernel-based PPI and QPI particularly struggle to reach the nominal level for Models 2 and 3.
5 Real Data Analysis
The real-data analysis uses moving-window prediction on selected M4 yearly series to compare DNN- and kernel-based prediction intervals across window sizes and model orders. Across Tables 4 and 5, DNN-based intervals—especially PPI—are more stable, supporting PPI with DNN as the preferred real-data option.
- Evaluation design: The moving-window procedure trains models on each window, constructs j-step prediction intervals for future observations, and measures average coverage rate and interval length.Coverage is aggregated over n − J − w + 1 moving-window predictions for each forecast horizon.
- Experimental choices: The study uses simple first-order differencing, skips order selection, and instead compares multiple order values because order selection is non-trivial and the analysis focuses on pertinent DNN prediction.The comparison covers DNN and classical kernel estimators rather than preprocessing strategies.
- Results: DNN-based prediction intervals, especially PPI, are more stable than kernel-based intervals across window sizes and model orders, supporting PPI with DNN for real studies.Kernel intervals deteriorate when w/d is small, while their larger interval lengths and standard deviations when w/d is large create a practical drawback.
6 Discussion · A The Illustration of DNN Structure · B Proof
For additive time series models, DNN point predictions are consistent and prediction intervals asymptotically valid, while the proposed PPI improves finite-sample coverage. Simulations and real-data studies indicate greater dimensionality robustness than kernel estimators and better coverage than naive QPIs, alongside unresolved methodological and computational challenges.
- 6 Discussion: DNN point predictions are consistent and their prediction intervals are asymptotically valid under standard assumptions.These guarantees apply to time series generated by an additive model with a DNN estimator.
- 6 Discussion: The proposed pertinent prediction interval (PPI) is attainable under minimal assumptions and is designed to improve finite-sample prediction-interval coverage.The discussion presents an algorithm for constructing the PPI.
- 6 Discussion: The DNN-estimator prediction interval is more robust to the curse of dimensionality than the kernel-estimator variant.This comparative finding is supported by the simulation and real-data studies.
- 6 Discussion: The PPI achieves better coverage than the naive QPI in the simulation and real-data studies.The comparison concerns coverage rate rather than point-prediction accuracy.
- 6 Discussion: Practical order selection for kernel and DNN prediction remains unresolved, motivating a useful order-selection criterion.The discussion identifies order selection as an important direction for future work.
- 6 Discussion: Prediction inference for general models without the additive-structure constraint has not yet been explored.Extending the method beyond the additive model is identified as a future direction.
- 6 Discussion: Although pseudo-data DNN training can be parallelized, its memory demands remain heavy, motivating more computationally efficient PPI construction.The computational limitation concerns the cost of training DNNs on pseudo data.
- A The Illustration of DNN Structure: Figure 1 presents a simple fully connected DNN structure for reference.The supplied passage identifies the illustration but provides no further architectural details.
B.1 Proof of Theorem 3.1
The proof establishes the result first on [0,1]^d, constructs an approximating ReLU network, and then extends the argument to [−B_n,B_n]^d through domain rescaling. The construction embeds a sparse network into a fully connected equal-width architecture under a sample-size condition relative to its parameter count.
- The proof begins by analyzing the function class on the constrained domain [0,1]^d before extending it to [−B_n,B_n]^d.
- Using Yarotsky and Zhevnerchuk (2020), it constructs a general DNN with logarithmic depth in its parameter count and then embeds it into a fully connected equal-width DNN.The embedding uses Farrell et al. (2021) and preserves the constructed network's depth while controlling the required width.
- The network construction is feasible when the width is sufficiently large, with η ∈ (0,1), implying that the sample size exceeds the total number of network parameters.
- To handle the larger domain, the proof rescales [−B_n,B_n]^d to [0,1]^d using T(u)=2B_nu−B_n, transfers derivative bounds by the chain rule, and absorbs the rescaling into the DNN's first layer.
B.2 Proof of Theorem 3.2
The proof follows Brown (2024) while allowing ∥f0∥∞ to be bounded by a nondecreasing sequence {Dn}n∈N rather than 1. Localization and independent-block arguments preserve the proof details and yield the same error bound under assumptions B1–B4.
- B.2 Proof of Theorem 3.2: The proof is almost identical to Brown (2024), except ∥f0∥∞ is bounded by a nondecreasing sequence {Dn}n∈N rather than 1.The new bound does not alter the proof details in Brown (2024), Appendices D.2, D.5.2, D.5.3, and D.5.4.
- B.2 Proof of Theorem 3.2: The argument relies on localization analysis from Farrell et al. (2021) and independent blocks from Chen and Shen (1998).
- B.2 Proof of Theorem 3.2: The same error bound holds for any sieve estimator and corresponding function class satisfying assumptions B1–B4.
B.3 Proof of Theorem 3.3 · B.4 Proof of Theorem 3.4
Theorem 3.3 is proved by verifying Theorem 3.2’s conditions under assumptions A1–A3 and the specified DNN architecture settings. Theorem 3.4 then builds on Theorem 3.3’s L2 consistency, using high-probability events and Lipschitz control to bound the estimator uniformly.
- B.3 Proof of Theorem 3.3: Theorem 3.3 follows by checking all four conditions B1–B4 of Theorem 3.2 for DNN estimators under assumptions A1–A3.The proof concludes that Theorem 3.2 applies after verifying the remaining conditions and derives the stated final rate.
- B.3 Proof of Theorem 3.3: The B1 verification uses Theorem 3.1 with Ln ≍ log(n) and Hn ≍ n^(d/(r+d))(1/2−KB) log(n), yielding the displayed approximation-rate expression.The proof rewrites the expression as O(n^(rKB−(r/(r+d))(1/2−KB))) and invokes the specified range of KB.
- B.3 Proof of Theorem 3.3: Condition B2 follows from the uniform bound supf∈Nn ∥f∥∞ ≤ 2Bn, while B3 uses Pdim(Fn) ≍ n^(2η) log^5(n) for logarithmic depth and polynomial width.For the Theorem 3.3 setting, the proof additionally notes that d/(d+p) < 1.
- B.3 Proof of Theorem 3.3: Condition B4 is checked with mn(Zt) := 2 supf∈Nn |Yt−f(Yt−1)|, using least-squares loss, stationarity, and bounds implied by the estimator class.The proof controls the indicator 1{mn(Zt)≥8Bn} through 1{Yt≥Bn} and uses mn(Zt) ≤ 2(|Yt|+2Bn).
- B.3 Proof of Theorem 3.3: Theorem 3.3’s remaining probability and rate conditions are established for sufficiently large constants and sample sizes, allowing Theorem 3.2 to deliver the stated high-probability conclusion.The proof explicitly invokes the condition involving √n Bn a−˜ϵn(log n)(log log n) when Cδ is sufficiently large.
- B.4 Proof of Theorem 3.4: Theorem 3.4 starts from Theorem 3.3 by defining Jn := {∥f̂n−f0∥L2 ≤ ϵn} with P(Jn) → 1, then intersecting it with En to form Gn.Under C1, the proof states that the complement probability of Gn is controlled.
- B.4 Proof of Theorem 3.4: On Gn, the proof sets gn := f̂n−f0 and bounds Mgn := ∥gn∥∞ by exploiting Lipschitz continuity under C3 and C1.The argument relates the uniform norm to the L2 norm over the joint region formed by a ball and [−Bn,Bn]^d, then uses Ψn → 0 under C3.
B.5 Proof of Theorem 3.5
The proof follows Wu and Politis (2024a) but replaces their parametric mean-model argument with consistency of the non-parametric function estimator established by Theorem 3.4.
- The proof is mostly the same as Lemma 2.1 of Wu and Politis (2024a).
- Unlike Wu and Politis (2024a), this setting uses a non-parametric autoregressive mean model rather than a parametric one.
- The argument avoids their smoothness and Boldin (1983) conditions requiring parameter consistency because the full function f0 is consistently estimated by ˆf under Theorem 3.4.
B.6 Proof of Theorem 3.6
The proof establishes Theorem 3.6 by analyzing the two-step case J = 2, with higher-step and one-step cases following similarly. It shows the key error terms converge to zero in probability and derives the theorem’s final statement using Glivenko–Cantelli and triangle-inequality arguments.
- B.6 Proof of Theorem 3.6: The proof first reduces the conditional distribution comparison through the tower property and equivalent expressions involving the forward prediction map.Notation is simplified by setting f := f0 and introducing L(y, Ȳn, ϵn+1) for the relevant prediction residual.
- B.6 Proof of Theorem 3.6: The argument verifies the required intermediate limit by decomposing its left-hand side and analyzing the resulting terms separately.The first term is controlled under Theorem 3.5, while the second is shown to vanish using complementary convergence results and the proof’s regularity arguments.
- B.6 Proof of Theorem 3.6: The proof controls the estimator and innovation discrepancies using Taylor expansion, smoothness of f0, Lipschitz continuity, and Eq. (20), yielding convergence to zero in probability.The Taylor-expansion term converges under Theorem 3.4 when the smoothness order of f0 is at least 1, while the remaining term is handled through Lipschitz continuity and Eq. (20).
- B.6 Proof of Theorem 3.6: The J = 2 proof combines the componentwise convergence arguments to establish Eq. (5) of Theorem 3.6 in probability.The proof reduces the target expression to terms controlled by Theorems 3.4 and 3.5, then combines them to obtain convergence to zero.
- B.6 Proof of Theorem 3.6: The final statement, Eq. (6), follows by applying the Glivenko–Cantelli theorem together with the triangle inequality.The proof presents this as the final step after establishing Eq. (5).
B.7 Proof of Theorem 3.7
The proof establishes consistency of point predictions at arbitrary quantile levels using Lemma 1.2.1 of Politis et al. (1999), Theorem 3.6, and A5. It then derives asymptotic QPI validity and consistency of the optimal L2 point prediction under A6.
- Point predictions are consistent at any quantile level by applying Lemma 1.2.1 of Politis et al. (1999) with Theorem 3.6 and A5.
- The QPI is asymptotically valid because its two endpoints are point predictions at specific quantile levels.
- The optimal L2 point prediction is consistent, with the final equality in the proof following from A6.
B.8 Proof of Theorem 3.8
The proof establishes geometric ergodicity of the forward bootstrap series by verifying E1–E2, then proves convergence of its stationary distribution through six conditions adapted from Franke et al. (2002).
- Geometric ergodicity: The forward bootstrap series is geometrically ergodic because the fitted estimator satisfies E1 and the kernel-estimated residual density satisfies E2.E1 follows by taking C = 2B_n with some λ < 1; E2 is verified using the residual density estimator.
- Geometric ergodicity: Centering residuals in Algorithm 2 does not affect the asymptotic results, by Chebyshev’s inequality and boundedness of the fitted estimator and observations.
- Stationary-distribution convergence: The bootstrap stationary distribution converges in probability to the original stationary distribution after verifying six variants of Franke et al. (2002)’s conditions.The key conditions include a contraction bound, uniform estimator consistency on expanding sets, and vanishing probability outside those sets.