Source-linked AI summary
On the prediction of stationary functional time series
Alexander Aue, Diogo Dubart Norinho, Siegfried Hörmann
TL;DR
The paper addresses the limited general methodology and practical difficulty of predicting stationary functional time series beyond FAR(1) models. It reduces functional observations with FPCA, applies standard multivariate prediction to the score vectors, and reconstructs curve forecasts, with a functional FPE criterion selecting lag order and dimension automatically. Simulations and pollution-data application show predictions that are always competitive with and often superior to benchmark methods.
Problem
Prediction of stationary functional time series has largely focused on FAR(1) because general functional methodology is limited, while practical implementation remains difficult.
Method
The method uses FPCA for dimension reduction, fits multivariate time-series models to the score vectors, reconstructs curve predictors, and uses existing software and a functional FPE criterion.
Results
The proposed predictions are always competitive with and often superior to benchmark predictions in simulations and an application to pollution data.
Takeaways & Limitations
The approach provides an easily applicable prediction methodology that can automatically determine lag structure and model dimensionality.
Takeaways & Limitations
The paper assumes that the data are already available in functional form and does not address data preprocessing.
Abstract
from arXiv · showhide
This paper addresses the prediction of stationary functional time series. Existing contributions to this problem have largely focused on the special case of first-order functional autoregressive processes because of their technical tractability and the current lack of advanced functional time series methodology. It is shown here how standard multivariate prediction techniques can be utilized in this context. The connection between functional and multivariate predictions is made precise for the important case of vector and functional autoregressions. The proposed method is easy to implement, making use of existing statistical software packages, and may therefore be attractive to a broader, possibly non-academic, audience. Its practical applicability is enhanced through the introduction of a novel functional final prediction error model selection criterion that allows for an automatic determination of the lag structure and the dimensionality of the model. The usefulness of the proposed methodology is demonstrated in a simulation study and an application to environmental data, namely the prediction of daily pollution curves describing the concentration of particulate matter in ambient air. It is found that the proposed prediction method often significantly outperforms existing methods.
1 Introduction
Functional time-series prediction is technically difficult and has often centered on FAR(1) models, while practitioners face limited software and generalization options. The paper proposes an implementable FPCA-based reduction to multivariate prediction, with theory, model selection, and empirical evaluation addressing these gaps.
- Motivation and gap: General stationary functional prediction equations are difficult to solve and implement, contributing to research emphasis on technically tractable FAR(1) models.The paper identifies limited functional theory and estimation for more general cases as a continuing boundary.
- Motivation and gap: Functional prediction methodology remains difficult for practitioners because few ready-to-use software packages exist and manual implementation can restrict use beyond academia.The introduction names the far and ftsa R packages but says tailor-made procedures remain scarce.
- Proposed approach: The proposed algorithm applies FPCA, fits a vector time-series model to the resulting d-dimensional score vectors, and reconstructs the predicted curve through the Karhunen-Loève expansion.Existing functional-data and multivariate time-series software can support the reduction, score modeling, and reconstruction steps.
- Theory and model selection: The paper establishes a prediction-error bound implying asymptotic consistency and develops a functional final prediction error criterion that jointly selects lag order p and score dimension d.The criterion supports an automatic prediction process while the paper compares the approach with Bosq’s classical FAR(p) prediction.
- Proposed approach: Unlike approaches requiring a priori knowledge of the functional model, the method reduces dimension first and then permits a variety of existing vector-process tools, including further lags or exogenous covariates.This design avoids estimating functional operators directly and broadens the models that can be entertained.
- Methodological comparison: FPC scores being uncorrelated at lag zero does not ensure diagonal higher-lag autocovariances, so separate univariate score models may lose information in cross-score dependence.The paper states that this issue is demonstrated in a simulation study.
2 Methodology
The methodology predicts functional time series by reducing curves to FPC score vectors, applying standard multivariate forecasting, and reconstructing predicted curves. It supports flexible stationary processes and automatic selection of model order and dimension.
- Methodological scope: The approach avoids direct operator estimation and is not restricted to a particular functional time-series specification.This permits extensions beyond assumed FAR structures, including models based on multivariate innovations.
- Algorithm 1: The algorithm first computes empirical FPC scores from observed curves, retaining a chosen number d of principal components.The retained components can be selected by requiring a target fraction of explained variation.
- Algorithm 1: It then applies multivariate prediction methods to the score vectors for a fixed forecast horizon h.Available approaches include Durbin–Levinson, innovations, exponential smoothing, and nonparametric algorithms.
- Algorithm 1: The predicted scores are transformed back into functional predictions through a truncated Karhunen–Loève representation.The reconstruction combines predicted scores with the corresponding empirical eigenfunctions.
- Methodological scope: For FAR(p) processes, the method is motivated by the approximate VAR(p) behavior of the FPC score vectors.The paper states that this connection is made precise in the appendix.
- Model selection and implementation: The functional final prediction error criterion automatically selects both the lag order p and score dimension d.The implementation uses standard R packages for FPCA and VAR forecasting.
3 Predicting functional autoregressions
The paper compares the common FAR(1) operator-based predictor with score-vector VAR prediction and extends the analysis to higher-order functional autoregressions. It also analyzes dimension-reduction error and proposes data-driven joint selection of lag order and dimension.
- FAR(1) prediction: The common FAR(1) predictor estimates the autoregressive operator and applies it directly to the latest observed curve.It serves as a benchmark for the proposed methodology.
- FAR(1) prediction: Under the FAR(1) model, fitting a VAR(1) to FPC scores yields a predictor asymptotically equivalent to the standard operator-based predictor when d is the same.The result is stated under Assumption FAR and corresponding regularity conditions.
- Higher-order FAR models: For FAR(p), the process can be represented as a first-order autoregression in a product Hilbert space, enabling vector-functional predictors.The state vector contains the current and lagged functional observations, while the operator matrix contains identity, zero, and autoregressive operators.
- Higher-order FAR models: Theorem 3.2 establishes the theoretical score-based prediction result for FAR(p) processes with Hilbert–Schmidt autoregressive operators.The dimension-reduction error tends to the innovation variance as d increases.
- Dimension reduction: The dimension-reduction error reflects both unexplained variance from omitted principal components and their contribution to the operators’ Hilbert–Schmidt norms.The bound decreases with increasing d but does not itself provide a practical dimension choice.
- Model selection: The functional FPE criterion jointly chooses p and d, making the prediction procedure data driven and dependent on sample size.The paper reports excellent practical performance for this methodology.
4 Prediction with covariates
The covariate extension incorporates scalar, vector-valued, and functional exogenous variables into multivariate score-based forecasting. Functional covariates are reduced by FPCA before being combined with response scores.
- Covariate extension: The covariate procedure adds exogenous variables to predictions of a functional response time series.These variables may include lagged values of other functional series.
- Covariate representation: Response curves and functional covariates are separately transformed into empirical FPC score vectors, with potentially different retained dimensions.Scalar and vector-valued covariates can be used directly.
- Covariate representation: All covariate vectors are combined into a single vector R_e_n before multivariate prediction.The resulting procedure accommodates covariates defined on different spaces.
- Prediction: The h-step forecast uses response-score vectors together with covariate vectors in an appropriate multivariate algorithm, then reconstructs the predicted response curve.The one-step case can be characterized through projection equations involving response and covariate covariance matrices.
- Prediction: The covariate score model uses finite lag order m and coefficient matrices determined by covariance equations.In practice, the covariance matrices are replaced by sample versions.
- Practical constraint: Conditioning on all observed curves is avoided because the required high-lag covariance matrices cannot be reasonably estimated from the sample.The application fits a VARX(p) model and selects p and d using the adjusted fFPE criterion.
5 Additional options
The paper extends its prediction framework with alternatives for prediction beyond the basic FAR setting and provides a practical procedure for constructing uniform prediction bands. It also identifies estimation and modeling choices that affect computational efficiency and finite-sample accuracy.
- Additional prediction options: The framework can be extended beyond FAR models by treating a fitted FAR process as a best approximation to the underlying stationary functional time series.This extension uses an FPE-type criterion for model selection.
- Additional prediction options: The innovations algorithm offers a potentially more parsimonious alternative that can be updated quickly when new observations arrive.It is especially useful when observations arrive sequentially and prediction models need frequent updating.
- Prediction accuracy: Including too many lag values can reduce estimation accuracy because covariance estimates become less reliable for smaller samples and larger lags.The Innovations Algorithm estimates covariances Γ(k) over increasing lag values, creating this finite-sample trade-off.
- Prediction bands: The proposed prediction-band algorithm computes score-vector predictions, residual variability, and empirical bounds from out-of-sample residuals.The procedure uses residual standard deviations over t and selects constants so that a specified proportion of residuals satisfies the resulting bounds.
- Prediction bands: Prediction bands widen at intraday times with higher volatility because their width is scaled by the estimated residual variation function γ(t).The method does not impose an a priori restriction on γ, allowing it to reflect the data’s structure and variation.
- Prediction bands: The band method does not require particular model assumptions, and better in-sample performance between competing methods yields narrower prediction bands.Unreported simulation results indicate good finite-sample performance even for moderate sample sizes.
6 Simulations
The simulations evaluate the proposed method across several functional autoregressive and moving-average settings, including cases where cross-score dependence matters. Results show that modeling score vectors and selecting order and dimension data-adaptively can improve prediction and model selection relative to scalar or competing procedures.
- Simulation design: The proposed vector approach represents curves through finite-dimensional basis coefficients and applies multivariate prediction methods to the resulting score vectors.Functional operators are represented by D × D matrices acting on basis coefficients.
- Comparison with scalar prediction: When Ψ(1) generates the data, individual score sequences have little temporal correlation but first–third score cross-correlations remain substantial beyond lag 1.This dependence structure favors vector modeling over separate scalar score forecasts.
- Comparison with scalar prediction: When Ψ(2) generates the data, individual-score autocorrelations decay slowly while cross-correlations are zero at all lags, favoring the scalar method.The simulation therefore contrasts settings where cross-sectional dependence is either important or absent.
- Comparison with scalar prediction: For Ψ(1), vector-method MSE is less than half the scalar-method MSE in a majority of simulation cases, whereas scalar prediction is slightly favorable for Ψ(2).The comparison used estimated p and d for the proposed method but fixed p = 1 and d = 3 for scalar predictions.
- Comparison with standard functional prediction: The fFPE approach has slight advantages over BKR in almost all settings, including a 30%–40% lower MSE when κ1 = 0 and κ2 = 0.8.BKR almost always failed to select the correct order in that setting.
- Comparison with standard functional prediction: Mean squared prediction errors are relatively robust to dimension d but sensitive to order p, with order underestimation causing non-negligible increases.As sample size grows, MSEa approaches fFPE and the dimension selected by fFPE increases.
- Comparison with standard functional prediction: For longer FAR processes, fFPE selects average orders between p = 4 and p = 5, while BKR selects p = 0 more than 90% of the time.This highlights a substantial difference in order adaptation between the two procedures.
7 Predicting particulate matter concentrations
The application predicts daily PM10 concentration curves using functional principal-component representations and multivariate forecasting, with model choices selected automatically. Across rolling out-of-sample evaluations, the proposed methods generally improve prediction accuracy over the BKR benchmark.
- Data preparation: 175 daily PM10 curves were constructed from half-hourly observations using ten cubic B-spline basis functions.The data covered roughly one winter season in Graz, Austria, after preprocessing and weekly-seasonality adjustment.
- Functional representation: The first three functional principal components explain about 89% of variability in the pollution curves.The first component explains about 72%, while the second and third explain 10% and 7%, respectively.
- Functional representation: The first component represents an overall level shift, whereas the second captures an intraday trend and the third controls the prominence of diurnal peaks.These interpretations are obtained by perturbing the estimated mean curve with empirical eigenfunctions.
- Prediction evaluation: Prediction quality was assessed with five consecutive blocks, each producing 15 out-of-sample forecasts summarized by mean and median squared errors.The competing models also selected values of d and p for comparison.
- Prediction results: Except in the first period, the new method produced significantly smaller MSE and MED values than BKR.In the second and third periods, prediction errors were on average only about half as large as those from BKR.
- Prediction results: Including temperature-difference covariates through a two-dimensional exogenous regressor yielded a further significant improvement in mean and median out-of-sample squared prediction error.This covariate-adjusted method is called FPEX and uses FPCA for covariate dimension reduction.
8 Conclusions
The paper proposes an accessible functional forecasting methodology that reduces curves to principal-component scores, applies multivariate prediction, and reconstructs functional forecasts. Simulations and pollution-data results indicate that the method is competitive and often superior to benchmark predictions.
- 8 Conclusions: The methodology represents functional observations through functional principal components and predicts the resulting score vectors with existing multivariate methods.Functional predictions are reconstructed using a truncated Karhunen–Loève decomposition.
- 8 Conclusions: The approach is made rigorous for the predominant FAR(p) case while remaining easy to apply with existing software packages.The paper also describes extension to exogenous covariates.
- 8 Conclusions: Simulations and pollution-data applications suggest that the proposed method is always competitive with and often superior to benchmark predictions.The conclusion summarizes evidence from both the simulation study and the environmental application.
A Theoretical considerations
The theoretical appendix formalizes estimation properties for a class of functional time series and gives a FAR(p)-specific result under an additional contraction condition. These results support the asymptotic behavior of the empirical mean and covariance estimators.
- A Theoretical considerations: For a large class of functional time series, the empirical mean and covariance are stated to be √n-consistent estimators of their population counterparts.The appendix introduces a lemma making this statement precise for FAR(p) processes.
- A Theoretical considerations: Under the FAR assumptions and ∥(Ψ∗)^k0∥L < 1 for some k0 ≥ 1, the expected squared mean-estimation error satisfies E[∥µ̂n − µ∥2] = O(1/n).The result is stated as part (i) of Lemma A.1.
- A Theoretical considerations: The proof invokes L2- and L4-m-approximability results from prior work to establish the stated estimation properties.The cited approximation conditions differ between the corresponding parts of the lemma.
A.1 The VAR structure
The VAR structure arises by projecting functional autoregressions onto empirical principal components, producing finite-dimensional equations for score vectors and operator coefficients. The resulting regression formulation enables least-squares estimation, but its errors are nonstandard and dependent on past observations.
- A.1 The VAR structure: The projected equations can be rewritten as a linear regression and estimated by ordinary least squares in the VAR(1) case.The construction replaces population eigenfunctions and variables with empirical counterparts.
- A.1 The VAR structure: Projecting the FAR(1) equation onto empirical eigenfunctions yields equations for the functional principal-component scores.The projected innovations and score vectors are represented as finite-dimensional quantities.
- A.1 The VAR structure: The finite-dimensional coefficient matrix has entries given by inner products of the functional autoregression operator applied to eigenfunctions.The matrix encodes the operator in the selected principal-component basis.
- A.1 The VAR structure: Although the equations resemble VAR(1) equations, their errors are generally noncentered, dependent, and correlated with past observations.The errors also depend in a complex way on the preceding functional observation.
- A.1 The VAR structure: The score dimension d is random but fixed for a fixed sample size, and the derivation ignores these effects thereafter.This qualification accompanies the nonstandard regression formulation.
A.2 Proof of Theorem 3.1
The proof establishes the asymptotic equivalence of two predictors by expressing both in finite-dimensional coordinates and controlling differences between their covariance matrices. It uses stationarity, ergodicity, matrix-inverse identities, and eigenvalue convergence to complete the argument.
- Finite-dimensional representation: The finite-dimensional representations use empirical eigenfunctions, projected scores, covariance matrices, and the Kronecker product.The matrix entries are defined from projected observations and empirical eigenfunctions.
- Predictor comparison: The VAR(1)-based predictor and the alternative functional predictor differ formally only through their associated covariance matrices.Both predictors are represented using empirical eigenfunctions and finite-dimensional matrix formulations.
- Matrix comparison: The covariance-matrix difference is bounded by decomposing it through the difference between inverse covariance matrices and a sample cross-product matrix.The proof defines S and uses Δ=(Γ̂^-1−Γ̃^-1)S before bounding its norm.
- Matrix comparison: The inverse-matrix difference is controlled with the identity (A+B)^−1=A^−1−A^−1(I+BA^−1)^−1BA^−1 when the required inverses exist.The proof sets A=Γ̃ and B=Γ̂−Γ̃.
- Asymptotic control: Stationarity and ergodicity yield λ̂_d→λ_d almost surely, supporting convergence of the inverse covariance terms when λ_d>0.If λ_d=0, the model has lower dimension d′<d, and both estimators use at most d′ principal components.
- Conclusion: Combining the preceding estimates establishes Theorem 3.1.The proof concludes after collecting the comparison and convergence results.
A.3 Proof of Theorem 3.2
The proof of Theorem 3.2 represents the functional autoregression through a finite-dimensional vector autoregression with a decomposed error term. It then analyzes the best linear predictor using causality, orthogonality, and norm bounds.
- Vector representation: A functional autoregression of order p is represented as a d-dimensional vector autoregression with coefficient matrices Ψ_j and error vector E_n.The coefficient entries are defined through inner products of the autoregressive operators with basis functions.
- Error decomposition: The error vector decomposes as E_n=T_n+S_n, with coordinates involving projected innovations and omitted functional components.The projected innovation contribution is specified through inner products with the basis functions.
- Prediction: The best linear predictor of Y_{n+1} is characterized using observations Y_1,…,Y_n.The proof proceeds by estimating the contributions of the decomposed error terms.
- Prediction: Causality implies that the components in S_n and T_n are uncorrelated, simplifying the expectation used in the predictor analysis.This orthogonality permits separate treatment of the two error contributions.
- Norm bounds: Bessel’s inequality and Parseval’s identity provide bounds for the expected squared norm of S_n, while repeated Cauchy–Schwarz applications complete the estimates.The proof collects these bounds to finish the theorem.