Source-linked AI summary

Statistical Aspects of Wasserstein Distances

Victor M. Panaretos, Yoav Zemel

arXiv:1806.05500v3stat.ME

TL;DR

Wasserstein distances provide a geometrically meaningful way to compare probability distributions and have become tools for statistical theory, inference, and analysis in Wasserstein space. This review synthesizes their definitions, properties, applications, and computational aspects, while noting that computation can be difficult and some statistical rates may be slower than n^-1/2.

  • Problem

    Statistics needs a probability metric that captures distributional perturbations and domain geometry while supporting convergence, bounds, and inference.

  • Method

    The review presents the probabilistic and analytic formulations of Wasserstein distance and surveys their roles in asymptotic theory, inference, Wasserstein-space statistics, and computation.

  • Results

    Wasserstein distances support convergence and moment arguments, statistical inference, and analysis of random measures, with applications spanning empirical transport and Fréchet means.

  • Takeaways & Limitations

    Wasserstein distance is a versatile framework for comparing complex probability distributions and for developing statistical theory and methodology.

  • Takeaways & Limitations

    Quantisation and exact n-to-n-point transport can be computationally difficult, and empirical Fréchet-mean convergence can be slower than n^-1/2.

Abstract

from arXiv · show

Wasserstein distances are metrics on probability distributions inspired by the problem of optimal mass transportation. Roughly speaking, they measure the minimal effort required to reconfigure the probability mass of one distribution in order to recover the other distribution. They are ubiquitous in mathematics, with a long history that has seen them catalyse core developments in analysis, optimization, and probability. Beyond their intrinsic mathematical richness, they possess attractive features that make them a versatile tool for the statistician: they can be used to derive weak convergence and convergence of moments, and can be easily bounded; they are well-adapted to quantify a natural notion of perturbation of a probability distribution; and they seamlessly incorporate the geometry of the domain of the distributions in question, thus being useful for contrasting complex objects. Consequently, they frequently appear in the development of statistical theory and inferential methodology, and have recently become an object of inference in themselves. In this review, we provide a snapshot of the main concepts involved in Wasserstein distances and optimal transportation, and a succinct overview of some of their many statistical aspects.

1 Introduction

Wasserstein distances quantify optimal transportation between probability distributions and provide a geometrically meaningful framework for statistical theory and inference. This review introduces their probabilistic and analytic definitions, structural properties, and statistical roles.

  • Motivation: Wasserstein distances measure the minimal transportation cost between probability distributions and incorporate the geometry of their domain.They are proper metrics on measures with finite p-th moments.
  • Statistical roles: In statistics, they support asymptotic theory, methodological inference, and analysis in which probability measures themselves form the sample or parameter space.The review organizes these uses into three broad categories.
  • Definitions: The probabilistic definition minimizes E||X − Y||^p over all couplings whose marginals are the target measures.The equivalent analytic formulation minimizes the integrated transport cost over couplings in Γ(µ,ν).
  • Statistical properties: Wasserstein convergence combines convergence in distribution with convergence of p-th moments, making the metric useful for central-limit-theorem-type arguments and easy upper bounds.Any joint distribution with the correct marginals supplies an upper bound.
  • Definitions: A coupling assigns source mass to destinations while preserving both marginals; deterministic couplings instead transport each source location to one destination through a map T satisfying T#µ = ν.Under sufficient regularity, such deterministic couplings are optimal and unique when p > 1.
  • Optimal maps: When the source distribution has a density, the optimal transport map is characterized as the unique gradient of a convex function, whereas discrete sources may require nondeterministic couplings.This distinction reflects whether the relevant convex potential is differentiable on the support of the source measure.

2 Optimal Transport as a Technical Tool

Wasserstein distances serve as a technical tool for large-sample theory because their coupling structure supports bounds, convolution inequalities, and convergence arguments. The section connects these properties to Gaussian approximation, Markov-chain convergence, concentration, and comparisons with other probability metrics.

  • Basic properties: Wasserstein distances have scaling, translation, and product-measure properties that facilitate technical calculations in finite- and infinite-dimensional settings.The product relation is stated for p = 2, while the first properties follow from corresponding couplings.
  • Gaussian approximation: Subadditivity under convolution allows Wasserstein distance to quantify deviations from Gaussianity and supports central-limit-theorem arguments.The review describes coupling-based proofs and applications to asymptotic normality.
  • Gaussian approximation: For triangular arrays, uniform convergence in distribution together with uniformly controlled second moments yields convergence in W2 and joint asymptotic normality.The row length may diverge arbitrarily with n.
  • Markov chains: Wasserstein contraction estimates can establish exponential convergence of Markov chains to equilibrium.The contraction constant is related to the Wasserstein spectral gap.
  • Concentration and approximation: The W1 dual representation yields concentration inequalities for Lipschitz functions, while Wasserstein metrics also quantify approximation errors for point processes.These applications extend the technical role of Wasserstein distances beyond central-limit-theorem arguments.
  • Relations to other metrics: For the Prokhorov distance P, P^2(X,Y) ≤ W1(X,Y) ≤ (D + 1)P(X,Y), while on unbounded spaces W1 cannot generally be bounded by the bounded-Lipschitz metric.The displayed comparison assumes the setting in which D is finite.

3 Optimal Transport as a Tool for Inference

Wasserstein distances support inference through goodness-of-fit tests, asymptotic distributions, parameter estimation, and convergence-rate analysis. Their statistical behavior depends strongly on the sampling design, dimension, regularity, and whether the underlying measures are discrete or continuous.

  • Goodness-of-fit and testing: Wasserstein goodness-of-fit tests compare an empirical measure with a fixed distribution or with the closest member of a parametric family.The review covers one-sample, two-sample, dependent two-sample, location-scale, structural-relationship, and parametric-fit settings.
  • Asymptotic inference: Under alternatives, univariate Wasserstein statistics typically have sqrt(n)-scale normal limits, whereas null limits are typically n-scale and non-normal.The two-sample theory assumes m/n converges to a finite positive constant.
  • Asymptotic inference: Quantile-process asymptotics require regularity because inversion changes the Brownian-bridge limit to B(t)/F′(F^-1(t)).For weighted Wasserstein functionals, the limiting behavior differs according to finiteness of the relevant covariance integrals.
  • Asymptotic inference: For finitely supported measures, directional Hadamard differentiability yields generally non-Gaussian Wasserstein limits, while countably supported measures require an additional summability condition.Under alternatives, the finite-support statistic has a distributional limit at sqrt(n) scale.
  • Expected empirical distance bounds: Expected empirical Wasserstein convergence exhibits a dimension threshold: absolutely continuous measures have n^-1/d behavior when d > 2p, while discrete measures generally have n^-1/(2p) behavior.Separated support gives the discrete lower bound, and an absolutely continuous component produces the curse-of-dimensionality lower bound.
  • Expected empirical distance bounds: On the real line, compactly supported measures satisfy EW1 ≤ Cn^-1/2, while attaining analogous rates for p > 1 requires density and integrability conditions.The J1 condition is essentially a moment condition; Jp for p > 1 additionally entails smoothness and prevents the density from vanishing too quickly in the support interior.

4 Optimal Transport as the Object of Inference

This section shifts from using Wasserstein distances for statistical tasks to treating Wasserstein space itself as the sample space for inference on random measures.

  • The framework observes probability measures sampled from a random measure Λ in Wasserstein space and infers quantities about Λ’s law.

4.1 Fr´echet Means of Random Measures

Fréchet means provide a Wasserstein-space notion of averaging for random measures, addressing blurring that can arise from ordinary measure averages. Their existence, uniqueness, consistency, and limit theory depend on the geometry and regularity conditions.

  • Ordinary averaging can blur distinct point masses, whereas the Wasserstein average of two Dirac measures lies at their midpoint.
  • For p = 2, the population Fréchet mean is defined as a minimiser of the Fréchet functional, with an empirical counterpart based on observed measures.
  • Existence and uniqueness of Fréchet means are nontrivial and depend on the induced geometry; W2 has geometry close to Riemannian.
  • In W2(R^d), empirical Fréchet means always exist and are unique when at least one input measure is absolutely continuous; analogous population results require absolute continuity with positive probability.
  • A law of large numbers shows Fréchet means converge under second-level Wasserstein convergence of laws, including empirical laws from samples of random measures.

4.2 Fr´echet Means and Generative Models

The Wasserstein Fréchet mean is closely tied to deformation models: averaging warped random measures can recover an underlying intensity when the warp satisfies suitable mean and regularity conditions.

  • Choosing a metric and its Fréchet mean implicitly assumes a data-generating mechanism, and Wasserstein geometry is linked to warping or phase variation.
  • A random increasing homeomorphism T models time distortion in functional data, with the aim of recovering paths on an objective time scale.
  • If the warp has identity mean and the required convex-gradient structure, the underlying intensity λ is a Fréchet mean of the warped random measure Λ.
  • When λ is absolutely continuous and T sufficiently injective, the Fréchet mean is unique and equals λ.

4.3 Fr´echet Means and Multicouplings

Fréchet means in Wasserstein space are characterized through optimal multicouplings, linking joint transport plans to barycentres and extending the relationship beyond Euclidean spaces.

  • An optimal multicoupling of measures produces a joint random vector whose averaged law is a Fréchet mean.
  • When one measure is regular, transport maps from the Fréchet mean can construct the optimal multicoupling, making the two problems equivalent.
  • In complete separable barycentric metric spaces, Fréchet means correspond to laws arising from optimal multicouplings under a squared-distance cost.
  • For Dirac measures, the Wasserstein Fréchet mean is a Dirac measure at the Fréchet mean of the underlying points; analogous relations extend to p > 1.

4.4 Geometry of Wasserstein space

Wasserstein space can be locally linearised at a sufficiently regular Fréchet mean through optimal maps and a tangent function space. Its geometry is flat in one dimension and under compatible multivariate structures, but positive curvature complicates general multivariate means.

  • Tangent-space construction: At a unique absolutely continuous Fréchet mean λ, optimal maps provide the basis for approximating Wasserstein space by the tangent space.The tangent-space construction uses optimal maps from λ and assumes sufficient regularity.
  • Tangent-space construction: The tangent space Tan_λ lies inside L2(λ), whose norm is inherited from measurable vector-valued functions.An arbitrary measure can be represented through its optimal map relative to λ, with the identity subtraction centering the representation at λ.
  • Exponential and logarithmic maps: The exponential map sends a tangent vector r to (r + i)#λ, while the log map provides its inverse when λ is absolutely continuous.Exponential-map segments correspond to Wasserstein interpolants and geodesics.
  • Curvature and compatible measures: Wasserstein space is flat in one dimension, where quantile-map correspondence is an isometry and Fréchet means have an explicit barycentric form.The same explicit mean formula extends to compatible multivariate collections.
  • Curvature and compatible measures: Without compatibility restrictions, multivariate Wasserstein space is positively curved and Gaussian Fréchet means lack explicit covariance formulas when covariances do not commute.For commuting covariances, an explicit solution exists; curvature becomes unbounded near singular covariance matrices.

4.5 Fr´echet Means via Steepest Descent

Fréchet means in Wasserstein space can be computed by gradient-based iterations using logarithmic and exponential maps. The resulting method reduces a multitransport problem to successive pairwise transports, converging to the unique Gaussian mean and at least a stationary point more generally.

  • Gradient iteration: Steepest descent in Wasserstein space moves along the negative Fréchet-functional gradient using the exponential map.The gradient is expressed through the log map, as in differential-geometric optimisation.
  • Gradient iteration: The iteration reduces finding a Fréchet mean from a multitransport problem to a succession of simpler pairwise transport problems.This mirrors the decomposition used in generalised Procrustes analysis.
  • Convergence and optimality: In the Gaussian case, the algorithm converges to the unique Fréchet mean; more generally, it reaches at least a stationary point.The paper also gives an example showing that local minima need not be global.
  • Convergence and optimality: Smoothness and convexity of supports provide an optimality criterion under which a sufficiently smooth local minimum is roughly a global minimum.Smoothness alone cannot resolve the existence of nonglobal local minima.

4.6 Large Sample Statistical Theory in Wasserstein Space

Large-sample theory for random measures is well developed in one dimension and selected compatible or Gaussian settings, but remains incomplete in general Wasserstein spaces. Results include Gaussian-process limits, finite-support Gaussian limits, and evidence that convergence can be slower than n^-1/2.

  • Consistency and central limits: Consistency of empirical Fréchet means is a necessary foundation for statistical theory of random measures in Wasserstein space.The review identifies rates of convergence and central limit theorems as the next steps beyond general consistency.
  • Consistency and central limits: In one dimension, the empirical Fréchet mean, represented as an L2 map, converges after √n scaling to a zero-mean Gaussian process.The limiting covariance structure is that of the corresponding random tangent element.
  • Compatible multivariate settings: Compatible multivariate settings retain Hilbert-space embeddings that support extensions of these asymptotic results and enable principal component analysis.Convex PCA is presented as an alternative procedure.
  • Gaussian random measures: For finitely supported Gaussian random measures, √n-scaled empirical proportions have a Gaussian limit, and a delta method yields a central limit theorem for the covariance estimator.The result covers weighted Fréchet means and also addresses the case of two arbitrary measures.
  • Scope boundary: Beyond location-scatter settings, recent results suggest that empirical Fréchet means can converge at a rate slower than n^-1/2.This marks a boundary for transferring standard parametric-rate intuition to broader Wasserstein settings.

5 Computational Aspects

Computing Wasserstein distances and Fréchet means generally requires finite-dimensional approximations, linear programming, or iterative numerical schemes because closed forms are uncommon. Entropic regularisation improves computational complexity, while quantisation and discrete transport remain costly.

  • Discrete and continuous formulations: Outside one-dimensional and Gaussian cases, explicit Wasserstein distances and optimal couplings are rare.This motivates numerical approximations and optimisation-based algorithms.
  • Discrete and continuous formulations: For discrete measures, couplings are transport matrices whose entries specify transferred mass, and the Wasserstein cost is minimised under mass-preservation constraints.The resulting finite problem can be solved by linear programming.
  • Discrete and continuous formulations: The multimarginal Fréchet-mean problem becomes a linear program with a number of variables equal to the product of the support sizes.Alternative exact and polynomial-time approximate formulations use fewer variables or constraints.
  • Computational limitations: Quantisation permits exact finite linear programs but is difficult in practice, and n-to-n transport costs scale badly with n.Even one-dimensional measures rarely have explicit optimal quantisations.
  • Continuous and dual methods: Dynamic formulations approximate Wasserstein geodesics through a convex problem, while steepest descent operates in the dual variable.These methods trade the optimal-map formulation for alternative optimisation structures.
  • Regularisation: Entropy regularisation makes the transport problem strictly convex with complexity n^2 instead of n^3 log n for linear programming.The regularised coupling is diffuse but converges to the sparse unregularised solution as κ approaches zero.

6 On Some Related Developments

Optimal transportation extends quantile concepts to multiple dimensions through maps from reference variables, supporting multivariate contours, depth, ranks, and regression. The review also notes that machine-learning applications of optimal transport are not covered due to space constraints.

  • Multivariate quantiles can be defined using an optimal transport map from a reference random variable, such as a uniform variable on the unit ball.
  • The resulting framework provides multivariate quantile contours and induces notions of depth and ranks that can be estimated from data.
  • Extensions address multivariate quantile methods without requiring finite variance and extend quantile regression to multivariate settings.
  • The review does not describe the fast-growing machine-learning literature on optimal transport because of space considerations.
Loading 1806.05500v3…