Source-linked AI summary

Riemannian Gaussian Distributions on the Space of Symmetric Positive Definite Matrices

Salem Said, Lionel Bombrun, Yannick Berthoumieu, Jonathan Manton

arXiv:1507.01760v2math.ST

TL;DR

The paper addresses the lack of a tractable probabilistic model for data in the SPD-matrix space P_m. It introduces Riemannian Gaussian distributions with exact normalisation, studies their inference, and derives an EM algorithm for mixtures; these tools connect maximum likelihood to Riemannian centres of mass and support mixture-based representation and classification. The results rely on the Rao-Fisher geometry and its negative-curvature structure.

  • Problem

    Existing literature lacked a rigorously defined and tractable probabilistic model that could represent statistical variability and support inference for data in P_m.

  • Method

    The paper introduces Riemannian Gaussian distributions, derives their exact normalising factor and sampling method, studies inference, and develops an EM algorithm for mixtures.

  • Results

    Maximum likelihood estimation of the Gaussian centre equals the Riemannian centre of mass, with additional asymptotic inference results for testing, covariance, and confidence regions.

  • Takeaways & Limitations

    Riemannian Gaussian mixtures provide a framework for representing and classifying SPD-matrix data within the Rao-Fisher geometry.

Abstract

from arXiv · show

Data which lie in the space $\mathcal{P}_{m\,}$, of $m \times m$ symmetric positive definite matrices, (sometimes called tensor data), play a fundamental role in applications including medical imaging, computer vision, and radar signal processing. An open challenge, for these applications, is to find a class of probability distributions, which is able to capture the statistical properties of data in $\mathcal{P}_{m\,}$, as they arise in real-world situations. The present paper meets this challenge by introducing Riemannian Gaussian distributions on $\mathcal{P}_{m\,}$. Distributions of this kind were first considered by Pennec in $2006$. However, the present paper gives an exact expression of their probability density function for the first time in existing literature. This leads to two original contributions. First, a detailed study of statistical inference for Riemannian Gaussian distributions, uncovering the connection between maximum likelihood estimation and the concept of Riemannian centre of mass, widely used in applications. Second, the derivation and implementation of an expectation-maximisation algorithm, for the estimation of mixtures of Riemannian Gaussian distributions. The paper applies this new algorithm, to the classification of data in $\mathcal{P}_{m\,}$, (concretely, to the problem of texture classification, in computer vision), showing that it yields significantly better performance, in comparison to recent approaches.

I. INTRODUCTION

The paper introduces Riemannian Gaussian distributions on the SPD-matrix space P_m to provide a tractable probabilistic model and statistical-inference framework. It gives exact distributional results, connects maximum likelihood with Riemannian centres of mass, and derives an EM algorithm for mixture estimation.

  • Motivation: P_m is an application-relevant space of symmetric positive definite matrices equipped with the Rao-Fisher metric and negative-curvature homogeneous-space structure.The metric is also called the affine-invariant metric and underlies the paper’s results.
  • Motivation: The paper addresses the lack of a rigorously defined and tractable probabilistic model for statistical variability in P_m.It introduces Riemannian Gaussian distributions for this purpose.
  • Riemannian Gaussian distributions: Riemannian Gaussian distributions depend on a centre parameter in P_m and a positive scale parameter, with density defined using Rao’s distance and Riemannian volume.The normalising factor plays the role of the Euclidean Gaussian normaliser.
  • Riemannian Gaussian distributions: An exact expression for the normalising factor, together with a sampling method, fully characterises these distributions theoretically and for simulation.Earlier treatments used asymptotic formulae valid only for small scale parameters.
  • Statistical inference: For samples from a Riemannian Gaussian, maximum likelihood estimation of the centre equals computing the Riemannian centre of mass.The estimate minimises the sum of squared Riemannian distances to the observations.
  • Statistical inference: The paper develops asymptotic and inferential results including hypothesis testing, asymptotic normality, covariance estimation, and confidence regions for the empirical centre.The asymptotic centre parameter is the unique global minimiser of the expected squared-distance functional.
  • Mixtures and scope: A new EM algorithm estimates mixtures of Riemannian Gaussian distributions for representing and classifying data in P_m.The mixture model is motivated by its expected richness for real-world SPD-matrix data.
  • Mixtures and scope: The statistical-inference and mixture results rely crucially on the Rao-Fisher geometry’s negative curvature and do not have meaningful counterparts under alternative metrics such as log-Euclidean geometry.The definitions and properties generalise immediately to other negatively curved Riemannian homogeneous spaces, but the paper focuses on P_m.

II. RIEMANNIAN GEOMETRY OF COVARIANCE MATRICES

The section defines the SPD-matrix space P_m and equips it with the Rao-Fisher metric. It develops the associated distance, volume, geodesics, and negative-curvature homogeneous-space structure used later in the paper.

  • Definition: P_m is the space of real symmetric, strictly positive definite m × m matrices.Positive definiteness means Y† = Y and x†Yx > 0 for every x ∈ R^m.
  • Metric: The paper studies P_m under the Rao-Fisher metric.This metric is the geometric foundation for the subsequent constructions.
  • Geometric objects: The section gives formulas for the Rao-Fisher metric, Rao’s Riemannian distance, the volume element, and geodesic curves.These objects define the geometric quantities used in the distributional analysis.
  • Geometric structure: With the Rao-Fisher metric, P_m is a Riemannian homogeneous space of negative curvature.This property is fundamental to the later theoretical results.

A. The Rao-Fisher metric : distance, geodesics and volume

The Rao-Fisher metric measures lengths and distances on the SPD-matrix space, yielding unique geodesics and a corresponding volume element. Polar coordinates express these geometric quantities through eigenvalues and eigenvectors, with a correction for coordinate nonuniqueness.

  • Metric and distance: The Rao-Fisher metric measures the squared length of infinitesimal symmetric-matrix displacements.The metric is also written as ds2(Y) = ∥dY∥2.
  • Metric and distance: Rao’s distance is the infimum of curve lengths between two SPD matrices.Under the Rao-Fisher metric, the distance is realised by a unique geodesic because the space is complete, simply connected, and negatively curved.
  • Volume: The Riemannian volume element is derived from the Rao-Fisher metric and supports integration over Pm.The passage introduces volume as the geometric product of local lengths, with matrix-entry notation used in the resulting expression.
  • Polar coordinates: Polar coordinates parameterise each SPD matrix by eigenvalues and eigenvectors, but each matrix has m! 2m equivalent coordinate representations.Integrals therefore divide by m! 2m to correct for eigenvalue reorderings and eigenvector sign choices.

B. The Rao-Fisher metric : geometric properties

The Rao-Fisher geometry is homogeneous under GL(m) congruence transformations and matrix inversion, preserving its metric, distance, and volume integrals. Its negative curvature guarantees unique Riemannian centres of mass.

  • Homogeneity and invariance: GL(m) acts transitively on Pm by congruence transformations, making the Rao-Fisher space Riemannian homogeneous.For any Y and Z, an invertible matrix A can transform Y into Z while preserving the geometric structure.
  • Homogeneity and invariance: Congruence transformations and matrix inversion preserve the Rao-Fisher metric and distance.These isometries are stated through the invariance identities for ds2(Y) and d(Y, Z).
  • Volume invariance: Integrals with respect to the Riemannian volume element are invariant under the GL(m) action on Pm.This follows because the volume element is associated with a metric invariant under congruence transformations and inversion.
  • Riemannian centre of mass: A Riemannian centre of mass is a global minimiser of the variance function Eπ for a probability distribution π on Pm.The Rao-Fisher geometry guarantees that this centre exists uniquely and is the unique stationary point of Eπ.

III. RIEMANNIAN GAUSSIAN DISTRIBUTIONS

The paper formulates Riemannian Gaussian distributions on Pm with an exact density and develops their basic statistical inference. It links the distribution parameter to the Riemannian centre of mass and establishes asymptotic estimation results.

  • Definition and contribution: The paper’s main theoretical contribution is an exact formulation of Riemannian Gaussian distributions on Pm.The density is defined relative to the Riemannian volume element and depends on a centre parameter and scale parameter.
  • Statistical inference: The paper derives an exact normalising factor, sampling procedures, maximum-likelihood estimation, and hypothesis-testing methods for these distributions.These components are organised across the definition, sampling, and inference subsections.
  • Statistical inference: The parameter ¯Y of G(¯Y, σ) is the Riemannian centre of mass of the distribution.This connection underpins consistency and asymptotic normality results for its maximum-likelihood estimate.

A. Definition and basic properties

Riemannian Gaussian distributions use Rao’s distance and the Riemannian volume element, with a normalising factor independent of the centre. They admit transformation-based and polar-coordinate sampling schemes, including practical methods for evaluating or sampling the required densities.

  • Normalising factor: The normalising factor ζ(σ) is independent of the centre parameter ¯Y and has an exact expression involving the identity-centred distribution.Proposition 4 establishes ζ(¯Y, σ) = ζ(I, σ), while the latter is expressed using polar-coordinate integration and the multivariate Gamma function.
  • Normalising factor: For dimensions up to m = 50, the exact normalising expression can be evaluated using Monte Carlo integration.For m = 2, it also yields an analytic expression involving the error function.
  • Definition: A Gaussian distribution is defined relative to Rao’s distance and the Riemannian volume element, with parameters ¯Y ∈ Pm and σ > 0.The density notation f(Y | ¯Y, σ) is introduced using the volume element dv(Y).
  • Transformation properties: Congruence transformations map G(¯Y, σ) to G(¯Y · A, σ) for every A ∈ GL(m).This transformation property follows from the invariance of Rao’s distance and the Riemannian volume element.
  • Sampling: To sample G(I, σ), draw independent uniform U on O(m) and eigenvalue coordinates r from the specified joint density, then form Y = Y(r, U).Sampling G(¯Y, σ) additionally applies the congruence transformation by ¯Y 1/2.
  • Basic properties: If Y ∼ G(I, σ), then log det(Y) is normally distributed with mean 0 and variance m × σ2.This follows because log det(Y) equals the sum of the polar-coordinate variables r1, …, rm.
  • Sampling: When m = 2, sampling reduces to the univariate variables t = r1 + r2 and ρ = r1 − r2.The associated density factorises into separate densities for t and ρ.

B. Statistical inference problems

The section develops maximum-likelihood estimation and likelihood-ratio testing for Riemannian Gaussian distributions on Pm. Estimation identifies the location parameter with the empirical Riemannian centre of mass and determines dispersion through a unique equation.

  • Maximum likelihood estimation: Maximum-likelihood estimation of the location parameter equals the empirical Riemannian centre of mass of the samples.Both the location and dispersion estimates exist and are unique for every sample realisation.
  • Maximum likelihood estimation: The dispersion estimate is uniquely determined by the samples’ mean squared Riemannian distance from the estimated centre.The resulting function maps mean dispersion monotonically to the estimate of σ.
  • Hypothesis testing: The likelihood-ratio test compares dispersion around a specified value Y0 with dispersion around the estimated centre.The test rejects the null hypothesis when its statistic exceeds a threshold chosen for the desired significance level.
  • Hypothesis testing: The test statistic is more involved when σ is unknown because both hypotheses are composite in the nuisance parameter σ.If σ were known, the statistic would reduce to a comparison of the two dispersions.

C. Riemannian centre of mass and asymptotic properties

The paper establishes that the Gaussian parameter is the distribution’s Riemannian centre of mass and derives consistency and asymptotic normality for its empirical estimate. These results support inference about the centre, although asymptotic efficiency is left open.

  • Asymptotic properties: The asymptotic results make it possible to construct confidence regions and assign statistical significance to an empirical Riemannian centre of mass.The paper explicitly leaves the further property of asymptotic efficiency outside its scope.
  • Riemannian centre of mass: For any Riemannian Gaussian G(Ȳ, σ), the parameter Ȳ is its Riemannian centre of mass.This result underpins the subsequent consistency and asymptotic-normality results.
  • Asymptotic properties: The empirical centre ȲN converges almost surely to Ȳ as the sample size N tends to infinity.The proof combines convergence of empirical Riemannian centres of mass with the Gaussian distribution’s centre-of-mass identity.
  • Asymptotic properties: The scaled coordinate vector N^1/2(Δ1, …, Δp) is asymptotically centred normal with covariance 4σ^4 × C^-1.The coordinates represent the Riemannian logarithm of the empirical centre relative to Ȳ in an orthonormal basis.

IV. MIXTURES OF RIEMANNIAN GAUSSIAN DISTRIBUTIONS

The paper generalises Euclidean mixture modelling to Pm by introducing mixtures of Riemannian Gaussian distributions. It derives an EM estimator and a Bayes classification rule, motivated by the expected flexibility of sufficiently large mixtures.

  • Mixture modelling: Mixtures of Riemannian Gaussian distributions extend parameterised Euclidean mixtures to the Riemannian space Pm.Each component has a positive weight, centre Ȳμ ∈ Pm, and scale σμ > 0.
  • Mixture estimation: The paper derives a new EM algorithm for maximum-likelihood estimation of mixture weights, centres, and scales.The algorithm estimates ϑ = {(ϖμ, Ȳμ, σμ); μ = 1, …, M}.
  • Classification: A new Bayes classification rule uses the mixture representation to classify data in Pm.This replaces purely geometric nearest-neighbor classification rules used in recent methods.
  • Motivation and scope: The approach is motivated by the expectation that sufficiently large mixtures can approximate probability densities on Pm to arbitrary precision.The paper states that experimental verification of this expectation remains ongoing.

A. A new EM algorithm for mixture estimation

The paper develops an EM algorithm for estimating mixtures of Riemannian Gaussian distributions on the SPD-matrix space. It alternates conditional-expectation and maximization updates, with component centres updated as Riemannian centres of mass.

  • The algorithm generalises Euclidean mixture-estimation EM methods to the Riemannian geometry of Pm.
  • The number of mixture components is assumed known and fixed, so order selection is outside the paper’s scope.
  • EM iteratively updates mixture parameters through E steps computing expected complete log-likelihoods and M steps maximising them.
  • The model introduces latent membership labels, conditional membership probabilities, and conditional membership counts for each mixture component.
  • Each M step updates mixture weights, component centres, and dispersion parameters in sequence.
  • Each updated component centre is the unique Riemannian centre of mass associated with its conditional membership distribution.

B. Classification using mixtures of Gaussian distributions

The paper uses mixtures of Riemannian Gaussian distributions to cluster and classify SPD-matrix data. Its Bayes rule combines cluster size, dispersion, and Riemannian distance, extending nearest-neighbour classification.

  • Mixture models provide a framework for representing and classifying data in Pm beyond the usual Euclidean setting.
  • The proposed classifier uses posterior membership probabilities computed by the new Riemannian EM algorithm within a Bayes-optimal rule.
  • Supervised classification first clusters each labelled class, then assigns new observations to suitable clusters.
  • The approach uses mixtures of Gaussian distributions rather than Wishart mixtures to support classification based on Rao’s Riemannian distance.
  • The Bayes rule accounts for cluster size, dispersion, and distance from a test point to the cluster’s Riemannian centre of mass.
  • Ignoring cluster size and dispersion reduces the proposed rule to a nearest-neighbour rule based on Riemannian distance.

V. NUMERICAL EXPERIMENT

The numerical experiment evaluates three supervised classification rules on VisTex texture data. The proposed rule significantly outperforms nearest-neighbour and Wishart alternatives, while using three mixture components improves all rules.

  • The experiment applies the classifiers to texture classification, a problem involving predefined material or object classes.
  • The VisTex evaluation uses 40 images divided into 128 × 128-pixel patches, with 84 training and 85 classification patches per image.
  • The experiment compares the proposed rule with nearest-neighbour and Wishart classifiers on 85 × 40 classification patches.
  • The proposed classification rule provides significantly better performance than both the nearest-neighbour and Wishart rules.
  • For all three rules, training with M = 3 mixture components improves performance relative to M = 1.

APPENDIX A PROOF OF PROPOSITION 9

The appendix proves a proposition connecting the Riemannian Gaussian model’s energy function with its centre-of-mass characterization. The proof uses Riemannian gradients, Hessians, and differentiation under the integral.

  • The proof establishes that the Gaussian mean is a Riemannian centre of mass when it is a stationary point of the associated energy function.
  • The argument differentiates the probability-density normalization identity with respect to the Riemannian mean.
  • It uses the Riemannian gradient of squared Rao distance to express the gradient condition through logarithm maps.
  • The asymptotic argument uses a geodesic joining the population centre to the empirical centre and invokes the central limit theorem.
  • The proof relates the Hessian of the energy to a covariance bilinear form through the identity H = (1/2σ2) × C.
Loading 1507.01760v2…