Source-linked AI summary
On the choice of parameters in Singular Spectrum Analysis and related subspace-based methods
Nina Golyandina
TL;DR
The paper asks how SSA-related and subspace-based methods differ and how their parameters should be chosen across reconstruction, forecasting, and parameter-estimation problems. It analyzes these methods through a signal-plus-residual framework using theoretical results and simulations, finding that optimal choices and error behavior depend on the problem and residual structure. The main practical recommendation is that a window length near half the series length is often appropriate, with corrections for specific time-series classes.
Problem
The paper examines how to choose parameters for SSA-related and subspace-based methods across reconstruction, forecasting, and signal-parameter estimation.
Method
The paper uses a unified signal-plus-residual framework, theoretical analysis, and computer simulations to study errors, convergence, and parameter choices.
Results
Optimal parameter choices depend on the task and residual structure, while a window length near half the series length is appropriate in most cases.
Takeaways & Limitations
Window length should generally be chosen near half the series length, but corrected for time-series classes where the paper identifies different behavior.
Abstract
from arXiv · showhide
In the present paper we investigate methods related to both the Singular Spectrum Analysis (SSA) and subspace-based methods in signal processing. We describe common and specific features of these methods and consider different kinds of problems solved by them such as signal reconstruction, forecasting and parameter estimation. General recommendations on the choice of parameters to obtain minimal errors are provided. We demonstrate that the optimal choice depends on the particular problem. For the basic model `signal + residual' we show that the error behavior depends on the type of residuals, deterministic or stochastic, and whether the noise is white or red. The structure of errors and the convergence rate are also discussed. The analysis is based on known theoretical results and extensive computer simulations.
1 Introduction
The paper places SSA alongside related subspace-based methods, unifying their common features while distinguishing their applications to reconstruction, forecasting, and parameter estimation. It introduces Basic SSA and frames parameter choice and error analysis around signal-plus-residual time series.
- Scope and motivation: SSA-related methods share common algorithmic features but retain application-specific characteristics across signal processing and time-series analysis.The paper focuses mainly on SSA because it is not limited to a particular application area.
- Basic SSA: Basic SSA analyzes a time series as a sum of unknown identifiable components, such as a trend or regular oscillations.The method begins by embedding the observed series into a Hankel trajectory matrix using a chosen window length L.
- Basic SSA: SSA decomposes the Hankel trajectory matrix through SVD, producing eigentriples that are grouped to form component decompositions.Each eigentriple contains a singular value and corresponding left and right singular vectors.
- Basic SSA: Diagonal averaging reconstructs time-series components from the grouped trajectory-matrix components.Each reconstructed value is obtained by averaging along a secondary diagonal.
- Signal-plus-residual model: The framework treats an observed series as signal plus residual, where the residual may be random noise, a deterministic component, or a mixture.Useful reconstruction depends on approximate separability between the signal and residual; examples include trends, noise, and cyclic components.
- Problems and parameter choice: The paper studies reconstruction, forecasting, and parameter estimation, with parameter choice depending on the problem and the selected signal subspace.For finite-rank signals, subspace-based methods can estimate damping factors and frequencies from the leading eigenvectors.
2 Signal subspace
This section evaluates signal-subspace estimation errors in SSA using projector differences, emphasizing how window length and residual structure affect accuracy. Deterministic, white-noise, mixed, and red-noise residuals produce distinct error behaviors, including periodic improvements only in deterministic cases.
- Subspace error measure: Projector error is measured by the spectral norm between true and estimated signal-subspace projectors, equivalent to the sine of their largest principal angle.The estimated subspace is spanned by the leading left singular vectors of the observed trajectory matrix.
- Interpretation of accuracy: SNR alone does not determine SSA estimation error because it omits time-series length; even small-SNR sine waves can separate asymptotically from white noise.This asymptotic separation occurs as N →∞ with L ∼ βN, 0 < β < 1.
- Deterministic residuals: Deterministic periodic residuals yield periodic error behavior: when both L and K match component periods, projector perturbation becomes 0 through exact separability.Divisibility by only L or K produces one-sided orthogonality and generally reduces errors at special window lengths.
- White-noise residuals: White-noise residuals prevent exact orthogonality, so the error reductions observed at special window lengths do not occur.The comparison is shown on both logarithmic and original scales.
- Other residual structures: Mixed residuals inherit properties of both random and deterministic residuals, while red noise produces slightly worse accuracy than white noise at the same SNR.The simulations compare deterministic, random, combined, and non-white residual structures using MSD or RMSE.
3 Signal extraction
Signal-reconstruction error depends on window length, signal position, residual type, and alignment with periodic components. The paper recommends window lengths slightly below N/2, while deterministic residuals can make period divisibility more important than this adjustment.
- The reconstructed signal is obtained by diagonal averaging the reconstructed trajectory-matrix components.
- Signal-subspace estimation gives equivalent results for window lengths L and N − L + 1, while perturbation generally grows with increasing L.
- The pointwise optimal window length varies from N/3 to N/2, so the general recommendation is to choose L slightly below N/2.This choice particularly improves reconstruction at edge points relative to L = N/2.
- Window-length divisibility by the sine-wave period has a stronger effect on edge-point estimates than on estimates across the full series.
- For random residuals, the optimal window length is close to 0.4N; deterministic residuals can make divisibility by the sine-wave period more important.
4 Recurrent SSA forecast
Recurrent SSA forecasting constructs an LRF from the estimated signal subspace and applies it to reconstructed terminal data. Forecast accuracy is stable across many window lengths, but LRF-estimation and reconstruction errors respond differently to the reconstruction window.
- An LRF is a recurrence relation whose order is the number of coefficients, and the minimal LRF has the smallest governing order.
- Extraneous roots outside the unit circle can produce forecast terms that grow without bound when perturbed initial values are used.
- The min-norm LRF is obtained from the signal subspace and underlies recurrent forecasting from the last reconstructed signal points.
- LRF-estimation errors increase with window length and do not depend on Lrec, whereas reconstruction errors decrease as the window length increases.
- Forecast accuracy remains stable over a wide window-length range; LLRF and Lrec slightly below N/2 are appropriate, with small LLRF and Lrec near N/2 preferred.
- For certain deterministic residuals, even N and odd LLRF can make larger windows preferable because projector errors decrease as the window grows.
5 Subspace-based methods of parameter estimation
Subspace-based parameter estimation extends SSA by using estimated signal trajectories and shift invariance to recover characteristic roots. ESPRIT variants differ in estimation procedure, while TLS-ESPRIT is invariant to unitary changes of orthonormal basis.
- Subspace-based parameter estimation is presented as cohesive with SSA because both use signal-subspace information from trajectory matrices.
- A shift-invariant trajectory basis yields a matrix whose eigenvalues coincide with the roots of the signal’s characteristic polynomial.
- ESPRIT, HSVD, and HTLS estimate the shift matrix from an approximately shift-invariant estimated signal subspace.
- TLS-ESPRIT can be more accurate when the estimated signal subspace is approximate, because both sides of the shift relation contain errors.
- The TLS-ESPRIT estimate is unchanged by unitary transformations of the signal-subspace basis and is therefore the same for every orthonormal basis.
- For noisy-sinusoid frequency estimation, the first-order asymptotic variance has order 1/(K^2L) and is symmetric around N/2, implying optimal windows N/3 or 2N/3.
6 The rate of convergence
The paper compares convergence for fixed and N-proportional window lengths across deterministic, white-noise, red-noise, and combined residuals. Window choice and residual type strongly affect reconstruction, projection, and parameter-estimation errors.
- Window length L ∼ N/2: For signal-plus-constant series with L ∼ N/2, projector, whole-signal reconstruction, and frequency or exponential-rate estimates converge nearly at 1/N, 1/N^1.5, and 1/N^2, respectively.
- Residual type: Pure random residuals produce much worse convergence than deterministic constant residuals, while combined residuals inherit the worst case.
- Fixed window length L = L0: With fixed L, signal-plus-noise reconstruction does not converge, even when L is divisible by the signal or noise period.
- Fixed window length L = L0: With fixed L, signal-plus-constant reconstruction converges only when L or K is divisible by the period; otherwise, there is no convergence.
- Fixed window length L = L0: Small window lengths are generally unsuitable for reconstruction, and frequency estimation with small L converges only for pure independent white-noise residuals.
- Fixed window length L = L0: For red-noise residuals and small L, projection and frequency estimates may lack convergence even though their errors remain comparable in size to white-noise errors.The simulations indicate that accurate finite-length estimation can still occur despite nonconvergence; the reported projector-error magnitude is approximately 1.3·10^-3.
7 Choice of the window length and separability
Window-length choice for reconstruction depends on signal structure, separability, and available prior information. L close to N/2 is often effective, but shorter windows can be better for modulated or poorly separable signals.
- Modulated sinusoid: SSA-like methods can extract exponentially modulated sinusoids, extending beyond purely spectral analysis.Both damped and undamped sinusoids have SSA-rank 2 in the real-valued case.
- Modulated sinusoid: For slowly varying modulation, reconstruction can use L close to N/2 with four eigentriples or a short window covering roughly two periods with two components.The shorter-window approach treats the local amplitude as nearly constant, while modulation is represented in the right singular vectors.
- Modulated sinusoid: Simulations found no universally best window: L ∼N/2 performs better when approximate separability holds, whereas a couple of periods can be better otherwise.The shorter choice usually requires knowing the signal period, which is often unavailable.
- Modulated sinusoid: Under approximate separability, the recommended reconstruction uses L close to N/2 and retains eigentriples equal in number to the signal rank.This choice is accurate and does not require unknown period values, although a limited length range can reduce its adequacy.
- Separability: Complex modulation and weak separability can require considerably shorter windows because increasing L can increase approximating rank and reduce separability.Sequential SSA can first extract a trend with L1 = 30 and then extract periodicity with L2 = 100.
- Modulated sinusoid: For high noise, L near N/2 need not be optimal; L = 40 with two periods and two eigentriples slightly outperformed L = 200 with four eigentriples.The best choice can improve when additional information, such as the period, is available.
8 SSA processing of stationary time series
Toeplitz SSA and centering are useful for stationary series but can be misleading for non-stationary signals. The paper therefore distinguishes their stability benefits from their risks of bias, rank inflation, and incorrect forecasting.
- Toeplitz SSA: Toeplitz SSA can provide more stable reconstruction, forecasting, and estimates, but may be biased or inadequate for non-stationary series.The paper specifically warns against using it when trends or oscillations have changing amplitudes.
- Centering: Centering is less risky than using the Toeplitz covariance matrix, but it can either slightly improve or worsen SSA results and usually increases signal rank.For short sinusoidal series, centering can transform a rank-2 signal into rank 3 when the length is not divisible by the period.
- Centering: Single centering is intended for constant trends, whereas Double centering is described as suitable for linear trends.These variants are presented as alternatives to ordinary centering preprocessing.
- Non-stationary examples: The stationary-series recommendations are demonstrated on non-stationary examples, including exponential and sinusoidal signals.The paper frames these examples as tests of centering and Toeplitz SSA outside their usual stationary setting.
- Toeplitz SSA: Applying Toeplitz SSA to non-stationary finite-rank signals can increase the number of nonzero eigenvalues and require more eigentriples for accurate reconstruction.The resulting approximation may also have a wrong structure and produce an incorrect forecast.
9 SVD-origins of SSA and the choice of SSA parameters
The paper traces SSA parameter choices to several SVD-related traditions, including PCA, Hankel rank methods, spectral analysis, stochastic-process expansions, and dynamical systems. These origins emphasize different data structures and modeling assumptions.
- Principal Component Analysis: PCA-oriented SSA emphasizes manipulations of trajectory-matrix rows and columns, including centering and standardization.These operations shape interpretations of eigenvectors and factor vectors.
- Hankel rank-deficient matrices: Hankel rank-deficient methods connect time series governed by linear recurrent formulas with noisy signal processing and parameter estimation.Applications include damped or undamped cisoids and frequency estimation.
- Spectral analysis: Spectral analysis focuses on stationary time series, frequency characteristics, and testing for signals in red noise.This differs from SSA’s connection to singular values of trajectory matrices.
- Karhunen–Loève expansion: Karhunen–Loève expansion traditionally assumes zero expectation or subtracts known averages, with moving-average estimation extending the approach to trends.The estimated-average error is incorporated into the centered stochastic process.
- Dynamical systems: Dynamical-systems work contributed algorithms and parameter-selection ideas that helped establish SSA in several applied areas.The paper presents this as one origin among several rather than a single defining framework.
10 Conclusion
The paper unifies SSA-related methods through a signal-plus-residual formulation and studies their accuracy and parameter choices. Its simulations and error analysis support L close to half the series length as appropriate in most cases, while retaining problem dependence.
- Unified framework: A unified signal-plus-residual formulation reveals similarities and differences among SSA-related methods and their target problems.The framework covers the methods considered in the paper.
- Parameter choice: Error mechanisms and computer simulations support recommendations for choosing the SSA window length L.The recommendations are intended to minimize errors across representative settings.
- Parameter choice: L close to one-half of the time-series length is appropriate in most cases, but the optimal choice remains dependent on the problem.The conclusion presents this as a broad recommendation rather than a universal rule.