Source-linked AI summary
Registration of Functional Data Using Fisher-Rao Metric
Anuj Srivastava, Wei Wu, Sebastian Kurtek, Eric Klassen, J. S. Marron
TL;DR
Functional data can combine amplitude differences with phase differences caused by domain warping, complicating comparison and averaging. The paper uses the Fisher-Rao metric and SRVF representation to jointly align and compare functions through a quotient-space distance and Karcher mean template. The framework yields a consistent estimator under random warping, scaling, and vertical translation and performs better empirically than several recent alignment methods.
Problem
Functional observations may contain domain-warping variability that should be separated from vertical amplitude variability, while existing alignment approaches can decouple registration from comparison.
Method
The paper uses the Fisher-Rao Riemannian metric, whose warping invariance induces a quotient-space distance, together with SRVFs and a Karcher mean template for alignment.
Results
The framework provides a consistent estimator under random warping, scaling, and vertical translation and is empirically superior to several recently published functional-alignment methods.
Takeaways & Limitations
Separating phase and amplitude through a unified alignment-and-comparison framework supports analysis of functional data across growth, signature, neuroscience, and gene-expression applications.
Takeaways & Limitations
Some symmetric alternatives based on L2 alignment still do not define a proper distance on the function space.
Abstract
from arXiv · showhide
We introduce a novel geometric framework for separating the phase and the amplitude variability in functional data of the type frequently studied in growth curve analysis. This framework uses the Fisher-Rao Riemannian metric to derive a proper distance on the quotient space of functions modulo the time-warping group. A convenient square-root velocity function (SRVF) representation transforms the Fisher-Rao metric into the standard $\ltwo$ metric, simplifying the computations. This distance is then used to define a Karcher mean template and warp the individual functions to align them with the Karcher mean template. The strength of this framework is demonstrated by deriving a consistent estimator of a signal observed under random warping, scaling, and vertical translation. These ideas are demonstrated using both simulated and real data from different application domains: the Berkeley growth study, handwritten signature curves, neuroscience spike trains, and gene expression signals. The proposed method is empirically shown to be be superior in performance to several recently published methods for functional alignment.
1 Introduction
The paper addresses automatic functional-data alignment when domain warping mixes phase and amplitude variability. It proposes a unified geometric framework for registration and comparison, with applications including growth curves and signal estimation under random transformations.
- Motivation: Domain warping can mix variability in feature locations with variability in feature heights, motivating separate phase and amplitude analysis.The paper links this challenge to measurement uncertainty and inherent timing differences, such as maturity variation in growth curves.
- Applications: The framework is illustrated on growth data, signatures, neuroscience spike trains, gene expression signals, and simulations.The Berkeley growth study is used to examine common patterns after aligning individual growth curves.
- Motivation: Automatic unsupervised alignment is needed when temporal landmarks or natural warping models are unavailable.Manual landmark specification can be cumbersome for large datasets.
- Goals: The framework jointly performs warping and function comparison under a single objective rather than treating registration as preprocessing.This avoids using an alignment objective unrelated to the metric used for subsequent comparison.
- Goals: The paper also estimates a signal observed under random scaling, vertical translation, and time warping.The estimator is developed from observations generated from an underlying function under these random transformations.
- Proposed approach: The Fisher-Rao metric supplies a proper distance on the quotient space of functions modulo the warping group because it is invariant under domain warping.This distance supports computing a Karcher mean of function orbits and selecting an alignment template.
2 Function Representation and Metric
The paper represents absolutely continuous functions through square-root velocity functions and uses the Fisher–Rao metric to obtain computable, warping-invariant distances on function orbits.
- Function representation: Absolutely continuous functions are represented by square-root velocity functions (SRVFs), which are square-integrable and invertible up to vertical translation.Each SRVF corresponds to a function uniquely up to an additive constant.
- Fisher–Rao metric: The Fisher–Rao metric is invariant to domain warping, making it suitable for comparing functions whose parameterizations differ.The metric satisfies d_FR(f1, f2) = d_FR(f1 ◦ γ, f2 ◦ γ).
- Fisher–Rao metric: Under the SRVF representation, the Fisher–Rao Riemannian metric becomes the standard L2 metric.This transformation simplifies distance computation between functions.
- Elastic distance: The elastic distance compares SRVF orbits by minimizing the L2 distance over allowable warpings.The quotient-space distance is d([q1], [q2]) = infγ∈Γ ∥q1 − (q2, γ)∥.
- Elastic distance: This elastic distance is zero for domain-warped versions of the same function and is a proper distance between distinct orbits.It is a pseudo-distance on L2 itself but satisfies non-negativity, symmetry, and the triangle inequality on the quotient space.
3 Karcher Mean and Function Alignment
The framework computes a Karcher mean of SRVF orbits in a quotient space, selects a centered representative as a template, and aligns functions by elastic warping. Simulations show tighter alignment and reduced apparent amplitude variation, while the iterative cost decreases but need not reach a global minimum.
- Karcher mean: The method averages function orbits in the quotient space S, where time-warp-equivalent functions share one representation.It first computes the mean orbit [µ]n, then selects an element µn to serve as an L2 template.
- Karcher mean: The Fisher-Rao geometry of warping functions becomes spherical L2 geometry after representing each warp by the square root of its derivative.This representation maps warping functions to the unit sphere and permits Fisher-Rao distances to be computed as spherical arc lengths.
- Karcher mean: The Karcher mean is a local minimum of summed squared elastic distances, and its representative is centered when the associated warping functions have Karcher mean equal to the identity.Because the mean is an orbit, any rewarped representative is also a minimizer; centering selects a usable template.
- Function alignment: The alignment procedure iteratively estimates optimal warps, updates the mean, and returns a template, warping functions, and aligned functions.The update decreases the cost iteratively and converges, but convergence to a global minimum is not guaranteed.
- Simulation results: In simulations, alignment removes phase effects, sharpens common peaks and valleys, and narrows mean-standard-deviation bands while preserving amplitude differences.With no underlying phase variability, estimated warps remain near identity and aligned and original means are practically identical.
4 Signal Estimation and Estimator Consistency
The paper formulates signal recovery under random scaling, translation, and time warping, showing that complete alignment and Karcher means yield consistent estimators under stated assumptions.
- The observation model represents each function as a scaled, vertically translated, time-warped signal plus noise, with constant noise used for the consistency analysis.The analysis estimates the unknown signal or, equivalently, the warping functions from the observed functions.
- The Karcher mean of the observed SRVF orbits is the scaled orbit of the true signal SRVF, specifically ¯s[qg].This identifies the population template up to the orbit induced by time warping and an overall scale factor.
- The SRVF representation reduces the alignment problem to L2 optimization, using scale invariance and the isometric group action to identify optimal warpings.The supporting lemmas establish that positive scaling does not change the minimizing warping and that identity is optimal when comparing a function with its scaled version.
- If the inverse warpings have population Karcher mean γid and the signal derivative is nonzero almost everywhere, the center of the sample Karcher mean converges to E(¯s)qg.The result uses the estimated alignment warpings and their Karcher mean to select a representative from the mean orbit.
- Aligned functions recover the original signal asymptotically after correcting for the expected scale and translation, yielding E(¯c)g + E(¯e).The paper reconstructs g from aligned functions while assuming the population mean of the additive terms is known.
- In a simulation with n = 50 randomly warped, scaled, and translated sine signals, the estimated signal was reasonably successful despite large variability in the raw data.The example compares the true and estimated signals and evaluates estimation error across sample sizes n equal to 5, 10, 20, 30, and 40 using the L2 norm.
5 Experimental Evaluation of Function Alignment
The framework is applied to growth, signature, spike-train, and gene-expression data, producing visibly tighter alignments and enhanced peaks and valleys. An empirical comparison across simulated and real datasets evaluates five methods using three alignment criteria.
- Applications on real data: The experiments analyze functional data from growth curves, handwritten signatures, neuroscience spike trains, and gene-expression profiles.The study includes Berkeley growth data, signature acceleration functions, smoothed spike trains, and yeast cell-cycle expression profiles.
- Applications on real data: Alignment tightens the mean and variability of the Berkeley growth curves, revealing two average growth spurts between 3–4 and 10–12 years.The aligned mean shows enhanced peaks and valleys and suggests these two growth periods.
- Applications on real data: The signature and gene-expression examples show stronger peaks and valleys after alignment.For both datasets, the aligned functions and their cross-sectional means exhibit more exaggerated features than before alignment.
- Applications on real data: Smoothed spike trains from ten motor-cortex trials become well aligned, with more exaggerated peaks and valleys after registration.Gaussian-kernel smoothing with σ = 1ms converts spike trains into functional data before alignment.
- Comparisons with other methods: The comparison evaluates Fisher-Rao, AUTC, Tang–Müller, SMR, and MBM methods on three simulated and four real datasets using ls, pc, and sls.The three criteria assess variance, pairwise correlation, and derivative variance, with smaller ls and sls and larger pc indicating better alignment.
- Comparisons with other methods: Fisher-Rao performs uniformly well across the evaluation metrics, while sls appears to correlate best with visual alignment quality.The authors caution that ls, and sometimes all three criteria, can misrepresent alignment quality when signal shape preservation matters.
6 Discussion
The paper presents a parameter-free geometric approach for automated functional alignment and supports it with consistency results and broad empirical evaluation. It also identifies statistical modeling of phase variability as an important unresolved direction.
- Discussion: The method uses the Fisher-Rao metric and elastic distance between warping orbits to separate amplitude and phase variability.A Karcher mean orbit yields a template, with identity mean warping used as an additional selection condition.
- Discussion: The Karcher mean template is proved consistent for estimating a signal observed under random time warpings under basic conditions.The stated signal model also includes random scaling and vertical translation in the paper context, while the supplied discussion passage specifically states random time warpings.
- Discussion: Figure 12 evaluates five methods on three simulated and four real datasets using ls, pc, and sls, with best cases shown in boldface.The figure summarizes empirical alignment performance across all listed datasets and criteria.
- Discussion: Joint statistical models for amplitude and phase, especially phase modeling on the nonlinear manifold Γ, remain important future directions.The paper notes that functional principal component analysis is common for amplitude but cannot be directly applied to Γ.
A Proofs of Lemmas 1 and 2
These lemmas establish the differential behavior of the square-root mapping and the invariance of the transformed L2 geometry under warping.
- Lemma 1: The proof differentiates the square-root mapping Q for positive and negative inputs before combining the cases.The resulting directional derivative is expressed using the absolute value of the input.
- Lemma 1: Mapped tangent vectors are compared through their L2 inner product, linking the transformed representation to the original metric expression.The proof applies the derivative of Q to tangent vectors and evaluates their inner product.
- Lemma 2: The warping action preserves the L2 norm, which supports equality of distances after applying the same warp to both transformed functions.This norm preservation is used explicitly in the proof of the second lemma.
B Proofs of Lemma 3, Corollary 1, and Lemma 4
The proofs characterize optimal identity warping, establish uniqueness under a nonzero-measure condition, and show invariance of Fisher-Rao distance under common rewarping.
- Corollary 1: The identity warp minimizes the transformed distance between a function and its scaled version.The proof invokes the preceding lemma to reduce the minimization to the identity warp.
- Corollary 1: The minimizer is unique when the zero set of q has measure 0 because the cumulative squared integral F is strictly increasing.Strict monotonicity forces any minimizing warp to equal the identity.
- Lemma 4: Fisher-Rao distance is invariant when both warps are composed on the right with the same element of Γ.The proof attributes this to the isometry of the group action on Γ.
- Lemma 4: The optimal reparameterization can therefore be transferred through the common group action.The proof uses the distance invariance to relate the corresponding minimization problems.