Source-linked AI summary
Variational approach for learning Markov processes from time series data
Hao Wu, Frank Noé
TL;DR
The paper addresses how to learn useful feature mappings and Markovian models from time series when dynamics may be nonlinear, irreversible, or non-stationary. It introduces VAMP, based on Koopman-operator singular components and VAMP-r scores, and relates VAMP-2 to approximation error to construct VAMP-E for cross-validation. The reported results establish optimal model-parameter optimization through VAMP-r and optimal hyperparameter choice through VAMP-E cross-validation.
Problem
The paper addresses learning informative feature mappings and Markovian models from time series for dynamical processes that may be irreversible or non-stationary.
Method
VAMP uses Koopman-operator singular value decomposition to derive VAMP-r variational scores and relates VAMP-2 to approximation error to define VAMP-E for cross-validation.
Results
VAMP-r score maximization yields optimal model parameters, while VAMP-E cross-validation leads to an optimal choice of hyperparameters.
Takeaways & Limitations
VAMP provides a framework for optimizing feature mappings and Markovian models for arbitrary Markov processes, including irreversible and non-stationary processes.
Takeaways & Limitations
The major challenge in real-world applications of VAMP is overcoming the limitations of the resulting methods.
Abstract
from arXiv · showhide
Inference, prediction and control of complex dynamical systems from time series is important in many areas, including financial markets, power grid management, climate and weather modeling, or molecular dynamics. The analysis of such highly nonlinear dynamical systems is facilitated by the fact that we can often find a (generally nonlinear) transformation of the system coordinates to features in which the dynamics can be excellently approximated by a linear Markovian model. Moreover, the large number of system variables often change collectively on large time- and length-scales, facilitating a low-dimensional analysis in feature space. In this paper, we introduce a variational approach for Markov processes (VAMP) that allows us to find optimal feature mappings and optimal Markovian models of the dynamics from given time series data. The key insight is that the best linear model can be obtained from the top singular components of the Koopman operator. This leads to the definition of a family of score functions called VAMP-r which can be calculated from data, and can be employed to optimize a Markovian model. In addition, based on the relationship between the variational scores and approximation errors of Koopman operators, we propose a new VAMP-E score, which can be applied to cross-validation for hyper-parameter optimization and model selection in VAMP. VAMP is valid for both reversible and nonreversible processes and for stationary and non-stationary processes or realizations.
1 Introduction
The paper introduces VAMP to learn feature mappings and Markovian models from time-series data, addressing processes that may be nonlinear, irreversible, or non-stationary. It derives variational scores from Koopman-operator singular components and uses them for model optimization and cross-validation.
- Motivation: Time-series data from complex dynamical systems can often be transformed into features with approximately linear Markovian dynamics.Such models are easier to analyze, and collective changes can support low-dimensional feature-space descriptions.
- Motivation: The unresolved problem is choosing informative feature functions f and g at fixed dimension or with finite data.Minimizing regression error alone can select the uninformative constant model f(x) ≡ g(x) ≡ 1 and K = 1.
- Model selection: Variational scores can select the form and number of basis functions, unlike regression error, but finite datasets require cross-validation to avoid overfitting.Cross-validated variational scores can determine function classes, basis size, and other dynamical-model hyperparameters.
- Motivation: Existing variational approaches target dominant Koopman eigenfunctions, but their usefulness is limited for irreversible and non-stationary processes.Most real-world dynamical processes and their time series are described as irreversible and often non-stationary.
- VAMP framework: VAMP uses the singular value decomposition of the Koopman operator to identify optimal feature mappings for arbitrary Markov processes.The approximation error is minimized when f and g are the top left and right singular functions.
- VAMP framework: VAMP-r scores measure similarity to Koopman singular functions and optimize model parameters through maximization.The optimization is algorithmically identical to canonical correlation analysis on featurized time-lagged pairs and can use nonlinear approximators such as deep neural networks.
- Model selection: The VAMP-2 score is related to Koopman-operator approximation error, while VAMP-E supports cross-validation for hyperparameter selection.The paper reports that optimizing VAMP-E in a cross-validation framework leads to an optimal choice of hyperparameters.
2 Theory
The theory identifies optimal finite-rank linear Markov models with the top singular components of the Koopman operator and develops variational scores for estimating them from data. These results connect operator approximation errors to transition-density errors and apply broadly to Markov processes, subject to a Hilbert-Schmidt assumption.
- Koopman operator and finite-rank models: The Koopman matrix is the algebraic projection of the Koopman operator onto feature subspaces spanned by f and g.It predicts observables through E[h(x_t+τ)|x_t] = c^⊤K^⊤f(x_t) when h = c^⊤g.
- Optimal approximation: The optimal rank-k linear model uses the first k left and right singular functions and singular values of the Koopman operator.Specifically, f = (ψ1, ..., ψk)^⊤, g = (φ1, ..., φk)^⊤, and K = diag(σ1, ..., σk).
- Scope and limitation: The approach is universal for Markov processes but requires a Hilbert-Schmidt Koopman operator and is not applicable to deterministic systems in usual cases.The Hilbert-Schmidt condition supports the singular-value decomposition and finite-rank approximation.
- Example and interpretation: For a noisy one-dimensional system, the second singular functions indicate metastable states, while the third and fourth provide finer dynamical information.Combining the first four singular components yields a transition-density approximation with 6.6% relative Koopman-operator error.
- VAMP variational principle: VAMP variationally identifies the k dominant Koopman singular components by maximizing a family of VAMP-r scores.The principle permits any positive integer r and can be estimated from transition pairs in time-series data.
3 Estimation algorithms
The estimation algorithms represent feature functions with basis expansions and optimize their coefficients from time-lagged covariance matrices. In the resulting feature space, the best linear model is obtained with canonical correlation analysis.
- Feature representation: Feature functions f and g are represented as linear combinations of basis functions χ0 and χ1.Separate basis sets may be used, and the resulting feature functions can remain nonlinear in the original state variables.
- Feature representation: Matrices U and V combine m and m′ basis functions to approximate k Koopman singular components.U and V have sizes m × k and m′ × k, respectively.
- Data-driven estimation: The algorithms estimate covariance and time-lagged covariance matrices from trajectory data to optimize feature mappings and assess model quality.Multiple trajectories can be handled by averaging the corresponding covariance estimates.
- Feature TCCA: The same coefficient solution is obtained for any positive integer r in the VAMP-r score formulation.The VAMP-r score is represented in matrix form using the columns of U and V.
- Feature TCCA: Feature TCCA solves the best-linear-model problem by applying canonical correlation analysis to time-lagged data in feature space.The coefficient matrices are obtained through linear CCA on χ0(x_t) and χ1(x_t+τ).
2. Perform the truncated SVD
The truncated SVD extracts leading Koopman singular components from a normalized Koopman matrix and outputs a low-rank linear Markov model. Nonlinear TCCA can optimize the basis functions used for this approximation.
- Truncated SVD: The algorithm selects corresponding left and right singular vectors and converts them into singular values, singular functions, and a Koopman matrix.The resulting functions are fi = uᵀiχ0 and gi = vᵀiχ1, and the linear model is then output.
- Feature TCCA: Feature TCCA is equivalent to a least-squares regression model and generalizes EDMD to approximate Markov models for stochastic processes.When χ0 = χ1, the feature-TCCA formulation is identical to linear EDMD.
- VAMP-r scores: The maximal VAMP-r score for fixed basis parameters equals a Schatten-norm quantity involving singular values of the projected Koopman operator.The r-Schatten norm is the ℓr norm of the matrix singular values, with r = 2 giving the Frobenius norm.
- Nonlinear TCCA: Nonlinear TCCA optimizes basis-function parameters and can represent flexible structures including neural networks and decision trees.The optimized parameters determine the basis form; optimization can use gradient descent or other nonlinear methods.
- VAMP-E score: VAMP-E decomposes Koopman approximation error into an unknown constant and a model-dependent term that can be estimated from data.Maximizing VAMP-E is equivalent to maximizing R2 in feature TCCA or nonlinear TCCA.
4 Model validation
The paper uses cross-validation to select feature representations, model rank, and other hyperparameters while evaluating estimated Koopman models on held-out data. VAMP-E provides a stable validation score tied to approximation error.
- Cross-validation procedure: Cross-validation ranks different dynamical models by fitting each hyperparameter setting on training folds and scoring it on complementary test folds.The selected model or hyperparameter set maximizes the cross-validation score MCV(θ).
- Validation scores: Directly computing test-set VAMP-r scores is invalid because singular functions learned on training data are generally not orthonormal on test data.This mismatch motivates a validation score that does not require test-set orthonormality.
- VAMP-E validation: VAMP-E-based validation evaluates Koopman approximation error on test data and permits the model rank k to be fixed, maximal, or selected as a hyperparameter.This flexibility allows comparison of models with different numbers of singular components.
- VAMP-E validation: The VAMP-E validation score avoids matrix inverses and can therefore be computed stably.The paper contrasts this with possible numerical instability in subspace VAMP-r validation.
- Example 3: In Example 3, both cross-validation and exact VAMP-E scores peak at m = 33, although the training-set average continues increasing with m.Smaller m = 13 produces large approximation errors, whereas m = 250 causes overfitting and an estimated singular value above the true value.
- Lag-time selection: Lag time should be shorter than the timescale of interest and should satisfy approximate Chapman–Kolmogorov consistency across observables and multiples of the lag.The VAMP scores themselves cannot compare models with different lag times, so the Chapman–Kolmogorov test is used instead.
5 Numerical examples
Numerical examples show that nonlinear TCCA identifies meaningful slow or almost-invariant structures, predicts long-time distributions, and passes dynamical consistency tests in double-gyre and stochastic Lorenz systems.
- Double-gyre system: In the double-gyre system, the leading singular components are accurately estimated when five-fold VAMP-E cross-validation selects m = 37 basis functions.Smaller bases produce significant approximation errors, while larger bases are influenced by statistical noise.
- Double-gyre system: The estimated Koopman operator successfully predicts the long-time evolution of the double-gyre state distribution.The prediction is compared with the full simulation model from initial state (x0, y0) = (1.48, 0.8).
- Double-gyre system: The Chapman–Kolmogorov test confirms τ = 2 as a suitable lag-time choice for the double-gyre model.The test compares empirical and predicted time-lagged covariances, with uncertainty estimated from 100 bootstrap replicates.
- Stochastic Lorenz system: For the stochastic Lorenz system, singular-function patterns support coarse-graining into four macrostates associated with inner and outer basins of both attractor lobes.The sign boundary of ψ1 closely matches the boundary between almost-invariant Lorenz-flow sets.
- Stochastic Lorenz system: Projected leading singular components in a higher-dimensional observable space are almost identical to those computed directly from the original three-dimensional data.This illustrates the transformation invariance of VAMP.
6 Conclusion
The paper presents VAMP as a general framework for analyzing linearized Markov models across fields, with quantitative scores for evaluating and optimizing finite-data models. It also identifies high-dimensional optimization and probabilistic validity as important limitations and directions for future work.
- VAMP provides a general framework for analyzing linearized Markov models from different application communities.
- VAMP-r and VAMP-E scores quantitatively evaluate modeling accuracy.
- Feature TCCA, nonlinear TCCA, and VAMP-E-based cross-validation optimize models for finite dimensions and finite data sets.
- High-dimensional systems make the variational problem difficult to solve efficiently, motivating deep neural-network and tensor-decomposition approximations.
- Future work includes more general variational tensor methods and applications to metastable states, coherent sets, and dominant cycles.
- Operator-error optimization may produce transition densities with negative values, so valid probabilistic modeling is not guaranteed.
A Analysis of Koopman operators
The appendix analyzes Koopman operators underlying the variational framework, establishing approximation and estimation properties while specifying conditions for the theory. It also shows that deterministic Koopman operators may fall outside the compact Hilbert-Schmidt setting.
- For independent trajectories, covariance estimates used by the model are unbiased and consistent as S →∞.
- The approximate Koopman operator is the best rank-k approximation to Kτ in Hilbert-Schmidt norm.
- The leading singular components include (σ1, φ1, ψ1) = (1, 1, 1).
- The approximate transition density satisfies normalization but may be negative, so the resulting model need not be a valid probabilistic model.
- Sufficient conditions include finite state spaces or state space Rd with positive initial and lagged densities.
- For deterministic systems, the Koopman operator is not compact and therefore is not Hilbert-Schmidt.
C Proof of the variational principle
The appendix proves that the variational optimization recovers leading singular components of the Koopman operator or its feature-space projection. It also verifies feature TCCA as the corresponding computational solution.
- The variational optimization is equivalent to a constrained matrix problem whose optimum selects the leading singular directions.
- For reversible Markov processes, the maximum is achieved by eigenfunctions associated with the largest eigenvalues.
- Feature TCCA solves the stated optimization problem under orthonormality constraints.
- The projection formulation interprets the optimization as a variational problem for the feature-space Koopman operator.
- Ignoring statistical noise, feature TCCA returns exactly the singular components of the projected Koopman operator.
- The optimal estimation result is invariant to the choice of r ≥1.
E.3 An example of nonlinear TCCA
The nonlinear TCCA example evaluates VAMP-r on a stochastic system and describes decorrelation, optimization, and regularization procedures. The example also addresses numerical singularity and constrains the leading estimated singular value.
- The example uses a stochastic system driven by Gaussian white noise with mean zero and variance 1.
- The maximal R1 and R2 values occur at w = 0.3157 and w = 0.7069, respectively.
- PCA-based decorrelation transforms basis functions using empirical means and covariance matrices before nonlinear TCCA.
- The feature TCCA algorithm computes covariance matrices and outputs estimated singular components and feature functions.
- The estimated leading singular value satisfies σ1 ≤1, with 1 identified as the largest singular value of C.
- Nonlinear TCCA can suffer numerical singularity when covariance matrices are not full rank, motivating decorrelation or regularization.
H.2 Relationship between VAMP-2 and VAMP-E
The section relates VAMP-E to feature-based TCCA optimization and examines how subspace scores support model evaluation. It also identifies cross-validation limitations arising from hyper-parameter dependence and potentially singular covariance matrices.
- Feature TCCA maximizes VAMP-E by maximizing the first term on the right-hand side of equation (116).
- Nonlinear TCCA also maximizes VAMP-E because its first term vanishes while its second term is maximized over w.
- The subspace score measures consistency between feature-spanned subspaces and the dominant singular spaces, while remaining invariant under invertible linear feature transformations.
- Validation scores may be unreliable when the covariance-related matrices Vk are singular.
- Cross-validation cannot determine k when the validation score is independent of estimated singular components at k = max{dim(χ0), dim(χ1)}.