Source-linked AI summary

Equivalence of distance-based and RKHS-based statistics in hypothesis testing

Dino Sejdinovic, Bharath Sriperumbudur, Arthur Gretton, Kenji Fukumizu

arXiv:1207.6076v3stat.MEcs.LGmath.STstat.ML

TL;DR

The paper addresses the open connection between energy- and kernel-based statistics for two-sample and independence testing. It establishes an equivalence through distance-induced kernels and extends it to distance covariance, while examining consistency and test performance. The results show that alternative choices within the kernel family can yield more sensitive tests, subject to limitations in product-space distribution characterization.

  • Problem

    The link between energy distance-based tests and kernel-based tests remained an open question, including whether distance covariance has an energy-distance representation on a product space.

  • Method

    The paper constructs distance-induced kernels from negative-type semimetrics and uses RKHS equivalences to relate energy distance to MMD and distance covariance to kernel-based independence statistics.

  • Results

    The equivalence extends to distance covariance, and the paper reports that alternative distance-kernel choices can produce more sensitive two-sample and independence tests.

  • Takeaways & Limitations

    Energy distance and distance covariance belong to a broader family of kernel-based statistics that can be generalized to variables on general topological spaces.

  • Takeaways & Limitations

    For the product-space kernel, strong negative type characterizes independence but not equality of arbitrary probability measures on the product space.

Abstract

from arXiv · show

We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, maximum mean discrepancies (MMD), that is, distances between embeddings of distributions to reproducing kernel Hilbert spaces (RKHS), as established in machine learning. In the case where the energy distance is computed with a semimetric of negative type, a positive definite kernel, termed distance kernel, may be defined such that the MMD corresponds exactly to the energy distance. Conversely, for any positive definite kernel, we can interpret the MMD as energy distance with respect to some negative-type semimetric. This equivalence readily extends to distance covariance using kernels on the product space. We determine the class of probability distributions for which the test statistics are consistent against all alternatives. Finally, we investigate the performance of the family of distance kernels in two-sample and independence tests: we show in particular that the energy distance most commonly employed in statistics is just one member of a parametric family of kernels, and that other choices from this family can yield more powerful tests.

2.1. Semimetrics of negative type.

The paper generalizes energy distance to semimetric spaces of negative type and identifies conditions ensuring its expectations and nonnegativity. Negative-type structure also connects these spaces to Hilbert-space embeddings.

  • A semimetric is symmetric and separates distinct points by assigning distance zero only to identical points.
  • Negative type is characterized through a condition on finite point sets and real coefficients, yielding Hilbert-space representations.
  • All Euclidean spaces are of negative type, and ρ^1/2 is a metric of negative type whenever ρ is a semimetric of negative type.
  • The generalized energy distance is defined for probability measures on semimetric spaces of negative type, with moment conditions ensuring finite expectations.
  • Negative type guarantees that the energy distance is nonnegative, while strict positivity for distinct distributions is stated for the standard energy distance.

2.3. Distance covariance.

Distance covariance is extended from Euclidean variables to semimetric spaces of negative type. Its connection to energy distance is initially unclear, but the paper resolves this through kernel methods.

  • The paper extends distance covariance from metric spaces to semimetric spaces of negative type.
  • Generalized distance covariance measures dependence between random variables using semimetric-based expectations under finite-moment conditions.
  • Distance covariance characterizes independence when the underlying metrics have strong negative type.
  • The paper asks whether distance covariance equals an energy distance on the product space, since the direct product semimetric candidate is invalid.
  • The section motivates RKHS tools by introducing reproducing kernels and kernel embeddings for the later equivalence result.

3.2. Maximum mean discrepancy.

The paper formulates distribution comparison and dependence testing through RKHS embeddings. MMD compares embedded distributions, while HSIC applies MMD to a joint distribution and the product of its marginals.

  • MMD is the RKHS distance between embeddings of two probability measures and becomes a metric when the kernel is characteristic.
  • The product-space kernel uses component kernels on X and Y, with an RKHS isometrically isomorphic to their tensor product.
  • HSIC is defined as the MMD between the joint distribution and the product of its marginals.
  • HSIC also equals the squared Hilbert–Schmidt norm of the covariance operator between the component RKHSs.

4.1. Distance-induced kernels.

Negative-type semimetrics generate positive definite distance kernels through a chosen center. These kernels provide the RKHS representation needed to connect distance-based quantities with kernel statistics.

  • A semimetric generates a positive definite kernel by combining distances to a fixed center with the distance between two points.
  • The distance-induced kernel is defined from a negative-type semimetric and a center point.
  • Distance kernels are generally not strictly positive definite because their centered construction can produce nontrivial null directions.
  • The canonical RKHS feature map expresses the semimetric through Hilbert-space distances between embedded points.
  • For ℓ_q distances on subsets of R^d, the construction is valid for 0 < q ≤ 2 and yields a family of distance kernels.

4.2. Semimetrics generated by kernels.

The section characterizes how positive definite kernels generate negative-type semimetrics and how equivalent kernels share the same induced geometry. It also establishes moment conditions for kernel embeddings and the broader equivalence framework used for testing.

  • A nondegenerate kernel induces a valid semimetric of negative type, and kernels generating the same semimetric are called equivalent.
  • Distance kernels form only a subset of the kernels generating a given semimetric, so the kernel family is broader than distance-induced choices.
  • Equivalent kernels differ by additive shift functions while preserving the generated semimetric, although positive definiteness constrains valid shifts.
  • Finite half-moments with respect to the generated semimetric ensure that kernel embeddings and MMD are well defined, while stronger first-moment conditions ensure the energy distance exists.
  • The main results use this kernel–semimetric correspondence to establish equivalence between distance-based and RKHS-based two-sample and independence statistics.

5.1. Equivalence of MMD and energy distance.

This section proves that energy distance and MMD are the same statistic under the kernel–semimetric correspondence. It also describes the associated isometries and the moment conditions governing when each quantity is defined.

  • For any negative-type semimetric ρ and kernel k generating it, the MMD γ_k equals the corresponding energy distance when the required moments exist.
  • Equivalent kernels generate the same MMD, so the statistic depends on the shared semimetric rather than the particular kernel representation.
  • The maps z 7→k(·,z) and z 7→δ_z embed the semimetric space isometrically into the RKHS and probability-measure space, respectively.
  • For bounded kernels, the corresponding kernel embeddings are well defined for all probability measures in M1+(Z).
  • Finite half-moments suffice for MMD existence, whereas finite first moments are required for the energy distance and equality with MMD.

5.2. Equivalence between HSIC and distance covariance.

The section extends the kernel–semimetric equivalence from two-sample testing to independence testing. It identifies distance covariance as an HSIC instance and uses product-space kernels to obtain the correspondence.

  • Theorem 24 generalizes the correspondence to semimetrics of negative type, not only metrics, and derives it directly from the HSIC expansion.
  • Finite first moments of the marginals are imposed to ensure distance covariance exists, while corresponding kernel moment conditions suffice for HSIC.
  • Combining the two-sample and independence results gives a direct relation between energy distance and distance covariance.
  • Distance correlation can likewise be expressed using the associated kernels.

5.3. Characteristic function interpretation.

The section connects distance-based independence statistics with RKHS properties, showing how strong negative type corresponds to characteristic kernels while identifying limits of translation-invariant representations.

  • The characteristic-function approach uses translation-invariant kernels and Bochner representations for independence-related quantities.
  • The distance-covariance weight function is not integrable, so no continuous translation-invariant kernel exactly coincides with distance covariance.
  • The strong-negative-type condition therefore provides an RKHS criterion for whether distance-based statistics distinguish probability distributions.
  • A semimetric has strong negative type exactly when its associated kernel is characteristic to the relevant probability-measure space.
  • The product-space kernel can characterize independence without characterizing equality of all probability measures on the product space.

7.1. Two-sample testing.

The two-sample section develops empirical MMD/energy-distance tests under moment and operator conditions, using pairwise distances and asymptotic spectral distributions.

  • Empirical MMD estimates are well defined under a finite half-moment condition, matching the condition required for the corresponding energy distance.
  • When the kernel generates the semimetric, the empirical estimate uses only pairwise semimetric distances between sample points.
  • The centered kernel and its integral operator determine the null distribution of the degenerate two-sample V-statistic.
  • Trace-class assumptions ensure the relevant integral operator has the properties required by the asymptotic distribution theorem.

7.2. Independence testing.

The independence-testing section expresses HSIC through centered kernel matrices and derives its null distribution from products of kernel-operator eigenvalues under moment and independence assumptions.

  • HSIC is computed from Gram matrices for X and Y together with the centering matrix H.
  • Under the null hypothesis, the HSIC limiting distribution is a weighted sum of chi-squares.
  • The weights correspond to products of eigenvalues from the centered integral operators for the two marginal kernels.
  • The analysis requires the marginal distributions to satisfy the stated moment conditions so the associated integral operators are trace class.
  • The testing procedure estimates the null distribution using spectral methods based on empirical eigenvalues.

7.3. Test designs.

The paper develops spectral test designs for energy-distance and RKHS-based statistics, then evaluates distance-induced kernels across synthetic two-sample and independence problems. Results show that tuning the distance-kernel exponent can improve sensitivity, especially for fine-scale or high-frequency alternatives.

  • Test designs: Spectral estimation of test thresholds provides a cheaper alternative to bootstrap with indistinguishable performance, and extends to consistent finite-sample null distributions for HSIC.The approach uses empirical eigenvalues of centered Gram matrices for HSIC.
  • Experiments: The experiments compare distance-based and RKHS-based statistics on synthetic two-sample and independence data.Two-sample settings vary means, variances, and sinusoidal perturbation frequency; independence settings include rotated ICA data and sinusoidal dependence.
  • Two-sample experiments: When variances differ, the Gaussian kernel substantially outperforms the distance-induced kernel, while performance is similar when means or sinusoidal perturbations differ.The Gaussian-kernel advantage for variance differences decreases with dimension, where both methods perform poorly.
  • Two-sample experiments: For sinusoidal two-sample perturbations, q = 1/3 and smaller achieve virtually errorfree performance even at high frequencies, improving dramatically over q = 1 and the Gaussian kernel.Here q = 1 corresponds to the commonly used energy distance; smaller improvements also occur for differing means and variances.
  • Kernel choice: Higher distance-kernel exponents favor mean differences along one noisy dimension, whereas smaller exponents detect finer-scale distributional differences at higher frequencies.The exponent therefore controls sensitivity to the scale at which distributions differ.
  • Independence experiments: For independence testing, smaller exponents improve detection of high-frequency sinusoidal dependence, and spectral tests are more sensitive than the quadratic-form test on the ICA benchmark.The quadratic-form test failed to detect dependence on that dataset.
  • Conclusion: The paper concludes that energy distance and distance covariance belong to a broader class of RKHS discrepancy and dependence measures that can be selected to design more powerful tests.This framework also applies beyond Euclidean data when an appropriate negative-type semimetric or kernel is available.

APPENDIX A: DISTANCE CORRELATION

The appendix relates distance correlation to kernel dependence measures and shows that varying kernels yields a broad family applicable to multivariate and structured domains. It also distinguishes characteristic kernels from the narrower class of universal kernels.

  • Distance correlation: Distance covariance extends to distance variance and distance correlation, with distance correlation admitting a direct kernel interpretation.The kernel formulation connects dependence measurement to RKHS covariance operators and Hilbert–Schmidt norms.
  • Distance correlation: Kernel-based distance correlation is invariant to joint scaling for homogeneous semimetrics and to translations for translation-invariant semimetrics.These invariances follow under the stated conditions on the semimetrics.
  • Applications: Varying the kernels produces a broad class of dependence measures for exploratory analysis on multivariate or structured, non-Euclidean domains.The construction generalizes the distance correlation of Székely, Rizzo, and Bakirov.
  • Universal and characteristic kernels: A universal kernel has an RKHS dense in the continuous functions on a compact metric space, and universal kernels are therefore characteristic.The stated characterization uses injectivity of the mean embedding on finite signed measures.
  • Universal and characteristic kernels: Distance kernels can be characteristic without being universal when they generate a semimetric of strong negative type; on compact subsets of R^d, the stated kernel is characteristic and nonuniversal for q < 2.The appendix notes that centered kernels equivalent to universal kernels retain characteristicness even after losing universality.
Loading 1207.6076v3…