Source-linked AI summary
Signature moments to characterize laws of stochastic processes
Ilya Chevyrev, Harald Oberhauser
TL;DR
The paper studies moment-based inference questions for probability measures on path spaces. It develops robust signature-based MMDs and signature kernels, showing that signature MMDs exploit sequence structure effectively in two-sample testing, while some classical comparisons are computationally costly.
Problem
The paper studies inference questions for probability measures when the underlying space X is a space of paths.
Method
The paper defines robust signature MMDs for path laws and discrete sequences, kernelizes them with signature kernels, and provides efficient recursive computation from finite samples.
Results
Signature MMDs generally outperform classical MMDs and efficiently exploit sequence structure, while their performance remains close when sequences have varying lengths.
Takeaways & Limitations
The authors recommend trying both linear and RBF signature MMDs across sample-size, sequence-length, and state-space regimes for two-sample testing.
Takeaways & Limitations
Classical comparison tests such as Hotelling and Friedman-Rafsky were reported only for short sequences and relatively low-dimensional state spaces because of their computational cost.
Abstract
from arXiv · showhide
The sequence of moments of a vector-valued random variable can characterize its law. We study the analogous problem for path-valued random variables, that is stochastic processes, by using so-called robust signature moments. This allows us to derive a metric of maximum mean discrepancy type for laws of stochastic processes and study the topology it induces on the space of laws of stochastic processes. This metric can be kernelized using the signature kernel which allows to efficiently compute it. As an application, we provide a non-parametric two-sample hypothesis test for laws of stochastic processes.
1. Introduction
The paper extends moment-based characterization from vector-valued variables to stochastic processes using robust signature features, then kernelizes them into a computable MMD for process laws.
- Motivation: The paper studies whether a feature map can support inference about functions and probability laws when the data space is a space of paths.The framework includes sequence-valued data through piecewise-linear interpolation and targets continuous-time processes.
- Limitations of ordinary moments: Signature moments generalize polynomial moments to path-valued data, but ordinary signature moments may fail to characterize even laws of processes with straight-line trajectories.Existing moment-decay conditions are difficult to verify and exclude processes such as geometric Brownian motion.
- Kernelized metric: The associated signature kernel induces a metric MMD on process laws, with topology comparable to weak convergence under natural assumptions.The kernel is formed as the inner product of robust signature features and supports finite-sample computation.
- Robust signatures: The robust signature uses a normalization map to provide a robust, universal, and characteristic feature map without the strong moment-decay assumptions otherwise needed.Robustness refers to a finite influence function, while characteristicness means expected features determine the process law.
- Application and scope: The robust signature MMD supports non-parametric two-sample testing and is designed to handle continuous-time paths, discretized sequences, and optional time-parameterization invariance.The normalization is essential for characteristicness and robustness; omitting it can make signature features non-characteristic.
2. Learning in Non-locally Compact Spaces
The section develops a feature-map framework for learning functions and probability measures on non-locally compact spaces. It combines bounded, point-separating features with strict-topology density to obtain universality and characteristicness, motivating tensor normalization for pathspace.
- Learning on a topological space is reduced to mapping data into a linear feature space and using linear methods for functions or probability measures.
- A feature map is universal when its linear functionals are dense in a target function class, and characteristic when integration against measures is injective.
- Universality and characteristicness are equivalent for locally convex topological vector spaces, allowing density arguments to establish measure separation.
- For non-locally compact spaces such as pathspace, choosing a function class whose dual contains probability measures and admits a density theorem is difficult, motivating bounded normalized monomial features.
- The strict topology on bounded continuous functions supports a Stone–Weierstrass-type result: point-separating subalgebras are dense, while its dual consists of finite regular Borel measures.
3. Robustification by Normalization
This section introduces tensor normalization, which bounds monomial-like features while preserving injectivity and algebraic structure. Applied to vector-valued data, it yields robust, universal, and characteristic features.
- 146
- A tensor normalization maps T1(V) injectively into a bounded subset through degree-wise dilation δλ(t)t.
- The normalization preserves boundedness, injectivity, and regularity under conditions on ψ, including boundedness, Lipschitz continuity, and ψ(x)/x^2 ≤ 1.
- Linear functionals of normalized features are B-robust, so their statistics have bounded influence even when the original monomials are unbounded.
- For vector-valued data, the normalized monomial feature map is a continuous injection, universal for Cb(Rd, R), and characteristic to finite signed Borel measures.
4. Robust Features for Smooth Paths
The section applies tensor normalization to signature features of absolutely continuous paths. The resulting features are robust, universal, and characteristic on both unparameterized and parameterized path spaces.
- 4.1 Iterated Integrals: Monomials of Paths: The signature consists of iterated integrals that generalize monomials from vector-valued data to paths, with shuffle products preserving algebraic closure.
- 4.1 Iterated Integrals: Monomials of Paths: The signature map is injective up to tree-like equivalence, which includes reparameterizations and can reduce dimensionality by ignoring time parametrization.
- 4.2 Features for Unparameterized Paths: Tensor-normalized signatures provide robust features for path laws while retaining the signature’s ability to distinguish paths up to tree-like equivalence.
- 4.2 Features for Unparameterized Paths: For unparameterized absolutely continuous paths, the normalized signature is universal for bounded continuous functions and characteristic to finite regular Borel measures.
- 4.3 Features for Parameterized Paths: Adding time as a state coordinate restores sensitivity to parametrization, yielding a genuine metric for parameterized paths and the corresponding normalized-feature results.
- 4.4 Features for Paths in General Spaces: A state-space map ϕ extends the construction to paths in general topological spaces, including evolving graphs when ϕ maps graphs to adjacency matrices.
5. Robust Features for Rough Paths
The rough-path extension enlarges robust signature features beyond absolutely continuous trajectories. It covers rougher paths and important stochastic-process classes while preserving universality, characteristicness, and robustness.
- The construction includes stochastic differential equations, semimartingales, Markov processes, and Gaussian processes whose trajectories may have unbounded 1-variation.
- Theorem 26 extends normalized signature features to geometric p-rough paths, whose larger p values permit rougher trajectories.
- For Pp, the normalized signature is a continuous injection into a bounded tensor subset, universal for bounded continuous functions, and characteristic to finite regular Borel measures.
- For every continuous linear functional, expectations of the normalized feature map define B-robust statistics on both rough-path and path spaces.
6. A Computable Metric for Laws of Stochastic Processes
The section constructs a robust-signature MMD for laws of stochastic processes, establishes its kernel and topological properties, and develops efficient computation for continuous paths and sampled sequences.
- 6. A Computable Metric for Laws of Stochastic Processes: The MMD is obtained by taking the supremum over the unit ball of the RKHS induced by the normalized robust signature kernel.Kernelization replaces the difficult supremum over a broad function class with an RKHS-based discrepancy that can be estimated from samples.
- 6.2 The Robust Signature Kernel and its MMD: The robust signature feature map yields a kernel whose MMD is a metric on finite signed measures over path spaces.The kernel is bounded, continuous, positive definite, universal, and characteristic in the stated settings.
- 6.2 The Robust Signature Kernel and its MMD: For finite-dimensional state spaces, weak convergence implies convergence in the robust-signature MMD, and the two topologies coincide on weakly compact sets.In general, convergence in the MMD does not imply weak convergence.
- 6.3 Lifting State Space Kernels to Pathspace Kernels: Lifting a state-space RKHS into a Hilbert-space path and applying the robust signature produces a bounded, continuous, positive definite pathspace kernel.The lifted kernel is characteristic and its MMD is a metric on the corresponding finite signed measures.
- 6.4 Computing the Signature Kernel: Discrete robust signatures and kernels operate on sampled sequences, including path-dependent partitions, while approximating their continuous-path counterparts with explicit convergence rates.Recursive signature algorithms provide computational bounds involving the number of time points, kernel-evaluation cost, and truncation level.
7. Application: Two-sample Tests for Stochastic Processes
The paper applies an MMD-based two-sample test to stochastic-process laws, comparing robust signature kernels with classical multivariate and kernel tests. Experiments show that signature MMDs exploit sequence structure effectively, while computational costs constrain classical statistics in higher dimensions.
- Test construction: The test evaluates H0: µ = ν against H1: µ ≠ ν using an unbiased MMD estimator combined with permutation testing.The permutation distribution provides significance thresholds, including the 95% quantile for a 5% test.
- Empirical results: The linear kernel MMD generally fails to distinguish different distributions, while RBF and Laplace MMDs outperform it but rarely match signature MMD power.Power improves with sequence length and sample size, while the qualitative ranking remains similar.
- Empirical results: Among classical statistics, Friedman-Rafsky performs well but is computationally expensive, while Hotelling is not competitive even for short sequences.These methods require minimal-spanning-tree construction or matrix inversion and were feasible only for short, low-dimensional sequences.
- Empirical results: Signature MMDs efficiently exploit sequence structure, whereas classical MMDs need much more data to reach similar test power.The comparison covers linear and nonlinear kernel MMDs alongside classical statistics.
- Robustness to sequence length: Signature MMDs remain applicable when sequences have different lengths, with only slight performance deterioration after randomly deleting sequence entries.Flattened-vector methods cannot directly handle varying sequence dimensions.
- Practical guidance: The authors recommend trying both linear and RBF signature MMDs across regimes involving small or large samples, sequence lengths, and state spaces.For the reported experiments, signature truncation levels were generally low, often at most 5 and frequently 2 for nonlinear kernels.
8. Summary
The paper extends classical moments to path-valued random variables through robust signature moments, yielding a universal and characteristic feature map that can be kernelized. The resulting framework covers continuous, discrete, stochastic, and structured sequential data while supporting non-parametric learning and testing of process laws.
- Robust signature moments generalize the classical moment map from finite-dimensional random variables to path-valued random variables.The resulting feature map is described as robust, universal, and characteristic, and can be kernelized.
- The framework encompasses discrete sequences, classical time series, stochastic differential equations, semimartingales, and processes evolving in space and time.Sequences can be embedded into paths, including through Donsker-type embeddings for linear state spaces.
- The approach also accommodates genuinely discrete data such as text and structured non-Euclidean objects when the state space has a kernel.Texts are represented as lattice paths in the free vector space spanned by letters.
- A universal and characteristic feature map for stochastic processes opens research directions in high-dimensional statistics and non-parametric learning of process laws.The paper specifically highlights non-parametric testing in high or infinite dimensions, where classical kernels can have decreasing test power.
- The construction relies on tensor algebras, whose graded tensor components support the iterated-integral representation used by signatures.The tensor algebra is presented as the free algebra containing the state space, with tensor convolution as its product.
B. Features for Geometric Rough Paths
The rough-path extension defines signatures for irregular paths by encoding their iterated integrals in a truncated tensor framework. It covers important stochastic-process classes, including semimartingales, Gaussian processes, and Markov processes.
- Rough paths extend iterated-integral signatures beyond finite-variation paths by encoding path increments in a truncated tensor algebra.A p-rough path is a continuous map with tensor-valued components satisfying the rough-path relations.
- A p-rough path lift associates a rough path with an underlying continuous path, and smooth paths admit canonical lifts.For p = 1 the construction recovers the bounded-variation setting, while C1 paths canonically define p-rough paths for every p ≥ 1.
- The construction uses p-variation geometry and time changes to handle paths defined on different time intervals.The relevant topology is induced using increasing bijections between time intervals, although the resulting topology can be non-Hausdorff.
- Rough-path topology identifies paths modulo time reparameterization and other tree-like equivalences.Tree-like equivalence includes reparameterizations and paths differing by back-tracking excursions.
B.2 Ordered Moments for Rough Paths
For geometric rough paths, the signature extends to all iterated-integral levels and characterizes paths up to tree-like equivalence. Quotienting by this equivalence yields the space of unparameterized rough paths embedded by the signature.
- The signature map extends to geometric p-rough paths, with iterated integrals canonically defined at every level.Its first floor(p) levels coincide with the rough-path components, while higher levels are supplied by the extended signature.
- Two geometric rough paths have the same signature exactly when they are tree-like equivalent.Thus the signature is injective after quotienting out tree-like equivalence.
- The space of unparameterized geometric p-rough paths is defined as the quotient by tree-like equivalence and equipped with the topology induced by the signature embedding.This construction parallels the bounded-variation setting while making the representation invariant to the identified path redundancies.
- Adding a time component makes the signature sensitive to the original parameterization while retaining the rough-path structure.The canonical lift of t ↦ (x1(t), t) embeds the original rough path into an augmented state space.
B.4 Topological Properties and MMD for Measures on Rough Paths
The rough-path quotient space is metrizable, and its signature-based MMD has nuanced relationships with weak convergence. The two topologies coincide on weakly compact finite-dimensional feature settings, but MMD convergence need not imply weak convergence globally.
- The rough-path quotient space is metrizable, and the associated metric space is complete; separability also holds when the underlying Banach space is separable.These properties follow from the signature embedding and continuity of canonical lifts of smooth paths.
- MMD convergence does not generally imply weak convergence for probability measures on rough-path space.There exist measures with dk convergence but without weak convergence.
- On a weakly compact set of probability measures with finite-dimensional feature space, dk and weak convergence induce the same topology.In particular, weak convergence implies convergence in dk on this restricted class.
- The topology mismatch arises because signature-feature convergence can occur for paths that do not converge in the rough-path topology.The construction uses sequences converging in a higher-variation rough-path space but not in the original one.
B.5 Example: Non-robust Signature Moments gone wrong
The example shows that ordinary signature moments can coincide for two path-valued random variables even when their laws differ.
- B.5 Example: Non-robust Signature Moments gone wrong: All signature moments of X and Y coincide although X and Y have different laws.The construction uses paths scaled by random vectors with independent lognormal and perturbed-lognormal components.
- B.5 Example: Non-robust Signature Moments gone wrong: The coincidence follows because the tensor moments of M and N are equal, so the expected iterated-integral tensors cannot distinguish their laws.The expected m-fold tensor integrals capture only the corresponding m-th moments of M and N.
C. Kernel Background
This section constructs the RKHS associated with a feature map and characterizes how universality and characteristicness transfer between the feature map and its kernel.
- C. Kernel Background: The Moore–Aronszajn theorem gives a unique Hilbert space with k as reproducing kernel, with H0 dense in that space.H0 is the span of the kernel sections kx.
- C. Kernel Background: A unique linear map Ψ identifies H0 with the span of feature vectors, and Ψ is injective and isometric onto its image.The construction follows from linear relations among kernel sections matching the corresponding relations among feature vectors.
- C. Kernel Background: The image of H0 in the dual feature space is dense in Ker(ı)⊥.This identifies the relevant dual subspace generated by the feature map.
- C. Kernel Background: Under continuous embedding of H0 into a locally convex space F, the induced map ı is continuous into F.The proof uses the decomposition of the dual space relative to Ker(ı).
- C. Kernel Background: Universality of Φ to F is equivalent to universality of k to F, while characteristicness of Φ to F′ is equivalent to characteristicness of k to F′.The equivalences follow from density and the relation between feature-map images and kernel sections.
C.1 Signature Kernel Discretization
The discretization section defines a kernel on sampled paths and establishes boundedness and positive semidefiniteness, including when partitions are random.
- C.1 Signature Kernel Discretization: M is a bounded, positive semidefinite kernel on X+.Its construction uses truncated tensor features and the signature-kernel framework.
- C.1 Signature Kernel Discretization: The discretized construction represents paths by sequences sampled along partitions π and π′.The sampled sequences xπ and yπ′ are formed from paths x and y over their respective time intervals.
- C.1 Signature Kernel Discretization: The same inequality remains valid when the partitions are random.The result extends the deterministic-partition argument through the representation of MMD as expectations of kernels.
- C.1 Signature Kernel Discretization: The proof establishes the result by combining boundedness of the underlying kernel with positive definiteness of the truncated tensor-space inner product.The argument also invokes Cauchy–Schwarz and the preceding lemmas.
- C.1 Signature Kernel Discretization: The truncation analysis uses factorial decay in the level M to control the approximation error.The estimate follows from bounds on tensor levels and technical lemmas for truncation.