Source-linked AI summary

Functional Data Analysis of Amplitude and Phase Variation

J. S. Marron, James O. Ramsay, Laura M. Sangalli, Anuj Srivastava

arXiv:1512.03216v1stat.ME

TL;DR

Phase variation—lateral displacement of functional features—complicates amplitude-focused analysis and can distort standard summaries and models. The paper surveys representations, objectives, and algorithms for separating phase and amplitude, using examples to compare approaches. It concludes that alignment should be tailored to the application: Fisher–Rao is most effective for peak alignment in spectral data, whereas excessive peak alignment can distract in growth curves.

  • Problem

    Phase variation complicates functional data analysis by distorting averages, inflating variance, degrading regression fits, and complicating the identification of amplitude versus phase.

  • Method

    The paper summarizes alternative representations, objective functions, statistical models, and algorithms for functional-data alignment and amplitude–phase separation.

  • Results

    Fisher–Rao is reported as most effective for peak alignment in spectral data, while too much peak alignment can distract in growth-curve data.

  • Takeaways & Limitations

    Alignment objectives and algorithms should be tailored to the application rather than applying a single notion of feature alignment universally.

  • Takeaways & Limitations

    Amplitude–phase separation is not generally identifiable because the distinction depends on prior knowledge, the application, and the goals of analysis.

Abstract

from arXiv · show

The abundance of functional observations in scientific endeavors has led to a significant development in tools for functional data analysis (FDA). This kind of data comes with several challenges: infinite-dimensionality of function spaces, observation noise, and so on. However, there is another interesting phenomena that creates problems in FDA. The functional data often comes with lateral displacements/deformations in curves, a phenomenon which is different from the height or amplitude variability and is termed phase variation. The presence of phase variability artificially often inflates data variance, blurs underlying data structures, and distorts principal components. While the separation and/or removal of phase from amplitude data is desirable, this is a difficult problem. In particular, a commonly used alignment procedure, based on minimizing the $\mathbb{L}^2$ norm between functions, does not provide satisfactory results. In this paper we motivate the importance of dealing with the phase variability and summarize several current ideas for separating phase and amplitude components. These approaches differ in the following: (1) the definition and mathematical representation of phase variability, (2) the objective functions that are used in functional data alignment, and (3) the algorithmic tools for solving estimation/optimization problems. We use simple examples to illustrate various approaches and to provide useful contrast between them.

1. INTRODUCTION

Functional data are curves or surfaces whose features can vary in height and timing, with lateral displacements termed phase variation. The paper motivates modeling these timing differences because they can distort summaries and analyses, then introduces time-warping functions and alignment questions.

  • Functional data represent experimental units distributed over lines and areas as curves and surfaces, with variation in height across the observed domain.
  • Wine spectra show that registration issues can make the mean curve differ substantially from the similar shapes of individual observations.The example includes 31 red, 7 white, and 2 rosé wines.
  • Phase variation denotes lateral displacement of curve features, distinct from amplitude variation in curve height.The paper illustrates phase variation through wine spectra, weather thresholds, and pubertal growth timing.
  • Time-warping functions map system time to clock time and are generally modeled as smooth, strictly increasing mappings with inverses for alignment.The paper notes that application-specific boundary behavior and differentiability may also constrain the warping functions.
  • Ignoring phase variation can inflate variances, degrade regression fits, and require additional principal components.Averaging the wine spectra also produces peaks that are lower and wider than most sample peaks.
  • The paper surveys alignment goals, amplitude–phase distinctions, optimization strategies, statistical models, and software, while identifying unresolved questions about phase data objects and feature definitions.

2. VIEWPOINTS AND GOALS

The paper frames amplitude and phase as context-dependent components whose separation can be represented through warping classes and model-based alignment. It emphasizes that either component, or their joint variation, may be the scientific focus.

  • 2.1 The Identification of Phase Variation: Amplitude and phase are not uniquely identifiable: the same linear-function variation can be assigned entirely to shifts, entirely to amplitude, or split between them.This ambiguity depends on prior knowledge about how variation is generated.
  • 2.2 Types of Phase Variations: Possible phase representations include uniform scaling, uniform shifts, affine transformations, diffeomorphisms, and warpings with flat regions.The preferred class depends on the application context.
  • 2.3 Some Goals for an Amplitude/Phase Analysis: Phase variation may be treated as a nuisance when amplitude is the target, as in wine spectra, crop-season budgets, or growth-curve shape.In these settings, timing is secondary to relative heights, accumulated resources, or shape characteristics.
  • 2.3 Some Goals for an Amplitude/Phase Analysis: Phase can instead be the primary information when event timing matters, including neuronal bursts, sowing windows, and maturation timing.Here, the time-warp functions become the center of attention.
  • 2.3 Some Goals for an Amplitude/Phase Analysis: Both amplitude and phase, including their joint variation, can matter when timing and magnitude are linked, as in pubertal growth.Early growth spurts are stronger and later spurts weaker, while adult final heights vary little with either factor.
  • 2.4 The Role of the Model in the Amplitude/Phase Partition: Model choice determines the amplitude/phase partition: transformations that improve fit to a proposed amplitude model are defined as phase.Candidate model spaces include means, templates, functional linear equations, principal components, and differential equations.
  • 2.5 Amplitude/Phase Separation via Equivalence Classes: Equivalence classes define amplitude by treating curves as equivalent when time warping transforms one into another, extending ideas from statistical shape analysis.A single curve’s time-warped versions form elements of its amplitude equivalence class.

3. SOME CURRENT CURVE REGISTRATION METHODS

The paper surveys registration methods for estimating time warpings, emphasizing that alignment requires carefully choosing both the approximation criterion and warping class. It contrasts landmark, dynamic-programming, L2-based, and Fisher–Rao/SRVF approaches, highlighting their practical trade-offs and examples.

  • Registration methods estimate warping functions by aligning observed curves with a template, but ordinary least-squares approximations are not generally viable.The approximation y(s) ≈ x0[h(t)] requires a carefully defined sense of congruence.
  • Dynamic Time Warping (DTW): Dynamic time warping uses dynamic programming on a fixed grid to find a globally optimal piecewise-linear warp, but may produce nonsmooth or unnecessarily warped regions.Regularization can address these smoothness and greediness problems.
  • Landmark Registration: Landmark registration interpolates warping functions from feature-time correspondences, yet landmarks may be invisible, ambiguous, laborious to record, and incomplete between features.Its evidence is discrete even though the warping function is intrinsically continuous.
  • Registration Using L2 Distance and Correlational Criteria: Correlation and minimum-eigenvalue criteria measure linearity between a template and a warped curve, defining alternative alignment objectives that can be applied to functions or derivatives.The appropriate loss function and warping class jointly define amplitude and phase variation.
  • The Square-Root-Velocity Function and the Fisher–Rao Metric: Fisher–Rao registration uses the SRVF representation and an invariant metric to compare amplitudes across equivalence classes rather than performing simple least-squares alignment of SRVFs.The formulation supports inverse consistency when the optimal warp is invertible.
  • Applications: Applications show that nonlinear SRVF warping can visibly improve growth-curve alignment, while registration reduced mean squared PCA residuals from 0.0052 to 0.0038.The latter corresponds to a squared multiple correlation of 0.47, indicating that phase modeling accommodates nearly half the variation around the unregistered fit.

5. DISCUSSION AND CONCLUSIONS

The paper reviews phase variability, the pitfalls of ignoring it, and approaches to phase–amplitude separation. It concludes that alignment remains application-dependent, requiring choices about warping, objectives, and algorithms.

  • The paper emphasizes phase variability and the pitfalls of ignoring it in functional-data analysis.
  • It surveys solutions to L2-based pinching, either by restricting warping or using alternative matching metrics.
  • Phase–amplitude separation remains unsolved because no single mathematical definition or algorithm works across most applications.
  • Alignment objectives should be tailored to the application: peak alignment can help spectral data but distract in growth curves.
  • The discussion focuses on real-valued functions while noting that the ideas can extend to higher-dimensional signals such as images.
Loading 1512.03216v1…