Source-linked AI summary

Hilbert space embeddings and metrics on probability measures

Bharath K. Sriperumbudur, Arthur Gretton, Kenji Fukumizu, Bernhard Schölkopf, Gert R. G. Lanckriet

arXiv:0907.5309v3stat.MLmath.ST

TL;DR

The paper addresses when the RKHS distance γ_k is a genuine metric and how its induced topology and finite-sample distinguishability behave. It characterizes suitable kernels, analyzes arbitrarily close distinct distributions, and relates γ_k to other probability metrics. The results provide conditions for metricity and weak-topology metrization, while showing that kernel scale and dimensionality affect practical discrimination and estimation.

  • Problem

    The paper asks which kernels make the RKHS pseudometric γ_k injective and how γ_k relates to other probability metrics and weak convergence.

  • Method

    The paper analyzes characteristic-kernel conditions, estimator behavior, close-distribution constructions, and topological comparisons for γ_k.

  • Results

    γ_k is a metric for broad kernel classes, can be arbitrarily small for distinct distributions, and metrizes weak topology under stated kernel conditions.

  • Takeaways & Limitations

    Using γ_k requires characteristic kernels and appropriate kernel choices because distinct distributions may remain difficult to distinguish from finite samples.

  • Takeaways & Limitations

    For alternative metrics such as Wasserstein and related classes, convergence rates can depend on dimension and become slow in high dimensions.

Abstract

from arXiv · show

A Hilbert space embedding for probability measures has recently been proposed, with applications including dimensionality reduction, homogeneity testing, and independence testing. This embedding represents any probability measure as a mean element in a reproducing kernel Hilbert space (RKHS). A pseudometric on the space of probability measures can be defined as the distance between distribution embeddings: we denote this as $γ_k$, indexed by the kernel function $k$ that defines the inner product in the RKHS. We present three theoretical properties of $γ_k$. First, we consider the question of determining the conditions on the kernel $k$ for which $γ_k$ is a metric: such $k$ are denoted {\em characteristic kernels}. Unlike pseudometrics, a metric is zero only when two distributions coincide, thus ensuring the RKHS embedding maps all distributions uniquely (i.e., the embedding is injective). While previously published conditions may apply only in restricted circumstances (e.g. on compact domains), and are difficult to check, our conditions are straightforward and intuitive: bounded continuous strictly positive definite kernels are characteristic. Alternatively, if a bounded continuous kernel is translation-invariant on $\bb{R}^d$, then it is characteristic if and only if the support of its Fourier transform is the entire $\bb{R}^d$. Second, we show that there exist distinct distributions that are arbitrarily close in $γ_k$. Third, to understand the nature of the topology induced by $γ_k$, we relate $γ_k$ to other popular metrics on probability measures, and present conditions on the kernel $k$ under which $γ_k$ metrizes the weak topology.

1. Introduction

The paper studies the RKHS-based pseudometric γ_k, focusing on when it becomes a metric, how close distinct distributions can be, and how its topology compares with other probability metrics.

  • Distances between probability measures support hypothesis testing, density estimation, limit theorems, and other statistical applications.
  • RKHS unit-ball function classes define γ_k, whose kernel-based form offers computational and domain-flexibility advantages over several alternative metrics.The framework applies to structured domains such as graphs and strings, while γ_k can be computed from kernel expectations.
  • γ_k has an mn/(m + n)-consistent estimator for bounded measurable kernels, with dimension-independent rates for translation-invariant kernels on R^d.The paper contrasts this convergence behavior with the potentially slow estimation of φ-divergences.
  • Characteristic kernels: The paper gives accessible characterizations of characteristic kernels, including integral strict positive definiteness and full Fourier support for bounded continuous translation-invariant kernels on R^d.It also extends the translation-invariant characterization to the d-torus and notes that compactly supported translation-invariant kernels on R^d are characteristic.
  • Dissimilar distributions with small γ_k: Even characteristic kernels may assign arbitrarily small γ_k distances to distinct distributions, especially when their differences occur at sufficiently high frequencies.This can make finite-sample discrimination difficult despite injectivity of the population embedding.
  • Weak topology: The paper shows that γ_k is weaker than several standard probability metrics and identifies kernel conditions under which it metrizes weak convergence.Universal kernels on compact spaces suffice, while sufficient conditions are also given for R^d.

2. Hilbert Space Embedding of Probability Measures

The paper represents probability measures through RKHS mean elements, making γ_k the distance between their embeddings. It develops equivalent representations that clarify computation and interpretation, including Fourier-domain and smoothed-density views.

  • RKHS embedding: Each probability measure P is embedded as a unique RKHS element λ_P representing integration against P.The embedding is obtained through the Riesz representation theorem for the integration functional T_P.
  • RKHS embedding: For bounded measurable kernels, γ_k(P,Q) equals the RKHS distance between the embeddings of P and Q for all probability measures.This removes the need to verify separately that distributions belong to the kernel-dependent measure class P_k.
  • Equivalent representations: The embedding distance has an equivalent expectation-based representation involving within-distribution and cross-distribution kernel expectations.This representation can be evaluated in closed form or by numerical integration, depending on k, P, and Q.
  • Equivalent representations: For bounded continuous translation-invariant kernels, γ_k is an L2-distance between characteristic functions weighted by the Fourier-transform measure of the kernel.Thus, the kernel determines which frequencies contribute to the distributional comparison.
  • Equivalent representations: With Lebesgue Fourier measure, γ_k becomes the L2-distance between densities in a limiting sense, while shrinking Gaussian bandwidths approach this distance.Analogous limiting behavior is stated for Laplacian, spline, and inverse multiquadratic kernels.
  • Equivalent representations: Under suitable convolution conditions, γ_k is proportional to the L2-distance between densities of independently perturbed variables X + N and Y + N.The perturbation interpretation compares smoothed versions of the distributions rather than their original densities.

3. Conditions for Characteristic Kernels

The paper gives practical characterizations of characteristic kernels, showing when the RKHS pseudometric becomes a metric on probability measures. It also derives domain-specific criteria, construction rules, and restricted-distribution results.

  • Integrally strictly positive definite kernels are characteristic on any topological space.This supplies a general sufficient condition for γ_k to be a metric on all Borel probability measures.
  • Translation-invariant kernels on R^d: For translation-invariant kernels on R^d, k is characteristic exactly when the Fourier-spectrum support supp(Λ) equals R^d.The condition is stated as easier to check than earlier density-based characterizations.
  • Translation-invariant kernels on R^d: For translation-invariant kernels on R^d, full spectral support also implies characteristicness despite isolated spectral zeros, as with B2n+1-splines.Gaussian, Laplacian, and admissible odd-order spline kernels all have spectrum support R.
  • Translation-invariant kernels on R^d: All compactly supported translation-invariant continuous bounded kernels on R^d are characteristic.This result also highlights a computational advantage of compactly supported kernels.
  • Kernel construction: Characteristic kernels can be constructed by adding any admissible k1 or multiplying by any nonzero k2 to an existing characteristic kernel k.Neither added or multiplied kernel is required to be characteristic.
  • Restricted distribution classes: Kernels whose Fourier-spectrum support is a proper subset of R^d can still be characteristic on P1 when that support has non-empty interior.Theorem 12 applies to compactly supported, absolutely continuous distributions with characteristic functions in L1(R^d) or L2(R^d).
  • Translation-invariant kernels on T^d: On the torus T^d, k is characteristic exactly when Aψ(0) ≥ 0 and Aψ(n) > 0 for every n ≠ 0.Equivalently, the Fourier-spectrum support is the entire Z^d; the Poisson kernel is characteristic, whereas Dirichlet, Fejér, and cosine kernels are not.

4. Dissimilar Distributions with Small γk

Characteristic kernels make γ_k a true metric, yet distinct distributions can remain arbitrarily close under it, especially when differences concentrate in high-frequency or high-order components.

  • The empirical estimate γ^2_{k,u}(m,m) is studied for B1-spline and Gaussian characteristic kernels across the constructed distributions.The figures vary ν and compare q = U[−1,1] with q = N(0,2).
  • For the B1-spline kernel, γ_k(P,Q) decays as |ν| increases and becomes harder to distinguish from zero with finite samples.The construction introduces increasingly high-frequency components as |ν| grows.
  • With Gaussian q, widening q can make the troughs in γ_k(P,Q) arbitrarily deep.The experiments compare uniform and Gaussian q for the B1-spline and Gaussian kernels using m = 1000 samples.
  • High-order kernel eigenfunctions can hide large distributional differences because their RKHS norms are large.A large |Pϕ_l − Qϕ_l| may contribute little to γ_k when l is large.
  • Characteristic kernels can still assign arbitrarily small γ_k to distinct distributions.Theorem 19 constructs P ≠ Q with γ_k(P,Q) < ε for any ε > 0.

5. Metrization of the Weak Topology

The paper compares γ_k with established probability metrics and identifies kernel conditions under which it metrizes weak convergence. These conditions differ between compact spaces and R^d.

  • Metric comparisons: A sequence may converge weakly while γ_k, Wasserstein, and total variation behave differently, showing γ_k can be weaker than W and TV.For the example sequence, W(P_n,P) = 1 and TV(P_n,P) = 1 while γ_k(P_n,P) tends to zero for several kernels.
  • Under bounded measurable kernels and the associated Hilbertian metric, γ_k is weaker than the Dudley, Wasserstein, and total variation metrics.The comparison holds without assuming that k is characteristic.
  • The comparison theorem can also provide bounds on Wasserstein and other metrics in terms of the computable γ_k.This is useful where closed-form expressions for β or W are unavailable or restricted.
  • Weak topology on compact spaces: On compact metric spaces, a universal kernel makes γ_k metrize the weak topology.In this setting, γ_k is equivalent to the Prohorov, Dudley, and Wasserstein metrics.
  • Weak topology on R^d: On R^d, γ_k metrizes the weak topology when k is translation-invariant with ψ ∈ C0(R^d) ∩ L1(R^d) and satisfies the stated sufficient condition.The result applies to kernels meeting Theorem 24’s assumptions.
  • Scope boundary: The entire Matérn class satisfies the R^d conditions, whereas Gaussian kernels do not.Characterizing suitable kernels on general non-compact domains remains open.

6. Conclusion and Discussion

The paper establishes conditions under which γ_k is a metric, shows that distinct distributions can nevertheless be arbitrarily close, and characterizes when γ_k induces weak convergence. It also discusses kernel-family extensions and the unresolved practical problem of choosing kernel parameters.

  • Integrally strictly positive definite kernels are characteristic, so their induced γ_k is a metric on probability measures.This condition is presented as natural and easier to verify than density-based universality conditions.
  • For bounded continuous translation-invariant kernels on R^d, γ_k is a metric exactly when the kernel’s Fourier-transform support is all of R^d.The characterization is explicitly checkable and also implies that compactly supported translation-invariant kernels are characteristic.
  • Distinct distributions can be arbitrarily close under γ_k, potentially making them difficult to distinguish from finite samples even when the kernel is characteristic.The paper presents this closeness result independently of whether the kernel is characteristic.
  • The paper relates γ_k to other probability-measure metrics and gives conditions under which it metrizes the weak topology.These results are intended to provide a theoretical foundation for applications in statistics and machine learning.
  • The paper leaves open how to choose a characteristic kernel parameter in applications, noting that inappropriate Gaussian bandwidths can make γ_k(P,Q) arbitrarily small.Thus, effective discrimination requires an appropriate parameter choice.
  • A supremum over a family of kernel distances, γ(P,Q)=sup{γ_k(P,Q): k∈K}, remains a pseudometric and is stronger than the corresponding Bayesian average α.If α distinguishes two distributions, γ also distinguishes them, but the converse need not hold.

Appendix A. Supplementary Results

The appendix collects analytical results supporting the paper’s main arguments, including convolution limits, Fourier transforms of measures, entire-function properties, and Paley–Wiener results.

  • Convolution with appropriately normalized and decaying functions converges to the original function at continuity points and almost everywhere under stated integrability conditions.Uniform convergence is given for bounded uniformly continuous functions, while the stronger decay assumption yields almost-everywhere convergence for L^r functions.
  • The Fourier transform of a finite Borel measure on R^d is introduced as a bounded, uniformly continuous function with additional stated properties.This result supplies the Fourier-analytic representation used in the paper’s kernel arguments.
  • The appendix cites the Riemann–Lebesgue lemma and a Paley–Wiener theorem for distributions as supplementary Fourier-analysis tools.It also cites a result stating that compact support makes the Fourier transform entire, and that an entire function vanishing on R^d is identically zero.
Loading 0907.5309v3…