Source-linked AI summary

An elementary introduction to information geometry

Frank Nielsen

arXiv:1808.08271v2cs.LGcs.ITstat.ML

TL;DR

Information geometry needs a coherent differential-geometric framework for studying information manifolds and their applications. This survey develops metric, connection, conjugate, and divergence-based structures, states the fundamental theorem, and applies them to information-science problems. It establishes dual curvature and flatness results while illustrating uses in decision problems and machine learning, with proofs omitted and some computational constructions approximated.

  • Problem

    Information geometry requires differential-geometric structures for information manifolds and their use in information-science problems.

  • Method

    The survey develops metric and connection frameworks, conjugate and statistical manifolds, α-manifolds, divergence-based connections, and applications in statistics and machine learning.

  • Results

    The fundamental theorem states that conjugate torsion-free connections have the same constant curvature, so one connection is flat if and only if its conjugate is flat.

  • Takeaways & Limitations

    Information manifolds provide structures for statistical decision rules, Bayesian hypothesis testing, and statistical-mixture clustering.

  • Takeaways & Limitations

    Proofs are omitted, Riemannian exponential mappings can be computationally intractable, and approximated Fisher information may be degenerate or structurally inaccurate.

Abstract

from arXiv · show

In this survey, we describe the fundamental differential-geometric structures of information manifolds, state the fundamental theorem of information geometry, and illustrate some use cases of these information manifolds in information sciences. The exposition is self-contained by concisely introducing the necessary concepts of differential geometry, but proofs are omitted for brevity.

1 Introduction

Information geometry studies information sciences through geometric structures, especially the geometry of decision making and model fitting. This survey introduces the required differential geometry, develops information-manifold structures, and illustrates applications in statistics and machine learning.

  • Information geometry geometrically investigates information sciences, including methods for distilling information from data into models.
  • A narrower view treats information geometry as the geometry of decision making, including model fitting as choosing parameters from a parametric model family.
  • Outline of the survey: The survey introduces manifolds equipped with a metric tensor and affine connection, then explains how this framework generalizes Riemannian manifolds.
  • Outline of the survey: It develops conjugate connections, statistical manifolds, α-manifolds, divergence-based constructions, dually flat manifolds, and exponential and mixture connections.
  • Applications: Applications include natural gradient descent, Bayesian hypothesis testing using Chernoff information, and clustering statistical mixtures in dually flat spaces.
  • Conclusion and appendix: The survey concludes with summaries, references, recent studies of principled distances and divergences, divergence estimation, and the canonical decomposition of multivariate Gaussian families.

2 Prerequisite: Basics of differential geometry

This section introduces manifolds as locally Euclidean spaces equipped with metric and connection structures. It develops how these structures define tangent-space measurements, vector transport, geodesics, curvature, and torsion.

  • 2.3.4 Curvature and torsion of a manifold: The section introduces intrinsic curvature and torsion as properties induced by the affine connection, alongside tensor and coordinate descriptions.Metric compatibility and conformality provide additional geometric structure for comparing tangent-space angles and lengths.
  • 2.1 Overview of differential geometry: Manifold (M, g, ∇): A smooth manifold is locally Euclidean and has tangent spaces that provide local linearizations for geometric reasoning.The tangent bundle collects all tangent spaces and has dimension 2D for a D-dimensional manifold.
  • 2.3 Affine connections ∇: Connections transport vectors between tangent spaces and define geodesics as autoparallel curves, extending ordinary Euclidean straightness.Because tangent spaces at different points are otherwise unrelated, a connection supplies the rule for comparing vectors along curves.
  • 2.2 Metric tensor fields g: Primal and reciprocal bases are mutually orthogonal, while their metric matrices satisfy G∗ = G−1.These bases support conversion between contravariant and covariant vector components and yield tangent-plane Mahalanobis distances.
  • 2.2 Metric tensor fields g: A metric tensor defines inner products on tangent spaces, allowing vector lengths, angles, and orthogonality to be measured.It is a smooth symmetric positive-definite bilinear form and induces a dual metric through reciprocal bases.
  • 2.3 Affine connections ∇: An affine connection is an independent differential-geometric structure that defines covariant derivatives, parallel transport, geodesics, curvature, and torsion.The covariant derivative differentiates one vector field with respect to another, while parallel transport associates vectors along smooth curves.

3 Information manifolds

Information manifolds equip a manifold with a metric and dual conjugate connections, yielding statistical, α-, and dually flat structures. These structures support curvature results, divergence-based constructions, dual Pythagorean identities, and projection uniqueness.

  • A conjugate connection manifold is (M, g, ∇, ∇∗), where the connections are conjugate with respect to metric g.
  • A statistical manifold is (M, g, C), with metric tensor g and totally symmetric cubic tensor C.
  • Any conjugate pair generates α-connections such that (∇−α, ∇α) are dually coupled, with ∇0 equal to Levi-Civita and ∇±1 equal to the original pair.
  • A torsion-free affine connection and its conjugate have the same constant curvature, and α-flatness occurs simultaneously for the dual α-connections.
  • A divergence induces an information manifold, while a strictly convex differentiable Bregman generator yields a dually flat manifold with dual affine coordinates and Pythagorean identities.
  • In dually flat spaces, projection uniqueness follows when the target submanifold is flat under the dual connection and the appropriate divergence is minimized.

3.8 Hessian α-geometry: (M, F, α) ≡(M, Fg, F∇−α, F∇α

Hessian α-geometry extends dually flat geometry by inducing dual affine connections from a convex potential function. At α = ±1, the resulting Hessian α-geometry is dually flat.

  • A dually flat manifold has a Hessian structure induced by a convex potential function F and supports a family of α-geometries.
  • At α = ±1, the Hessian α-geometry is dually flat.

3.9 Expected α-manifolds of a family of parametric probability distributions: (P, Pg, P∇−α, P∇α)

Expected α-manifolds build information-geometric structures from regular parametric probability families using statistical expectations. The section develops Fisher information for exponential and mixture families and illustrates the Cramér-Rao bound for normal distributions.

  • An expected manifold is built on a regular parametric family, with metric and cubic-tensor components expressed through statistical expectations.
  • The Fisher information matrix is positive definite for regular models and is invariant under sample-space reparameterization but covariant under parameter-space reparameterization.
  • For univariate normal distributions, the mean and standard deviation are orthogonal parameters, producing a diagonal Fisher information matrix.
  • 200 runs of 100 iid samples estimate normal parameters with MLEs; sample covariance ellipses differ from Fisher-information ellipses and their centers deviate from grid locations.
  • Exponential families and mixture families have Fisher information matrices that induce expected α-geometries with dual exponential and mixture connections.
  • Exponential and mixture families equipped with their dual connections form dually flat manifolds because the corresponding Christoffel symbols vanish.

3.10 Criteria for statistical invariance

This section asks which metrics, affine connections, and divergences are statistically invariant, and identifies the Fisher metric and f-divergences as central invariant structures.

  • Statistical invariance requires preserving the metric under important statistical mappings, including Markov embeddings and sample-space transformations.
  • The Fisher information metric is the unique invariant metric tensor under Markov embeddings, up to a scaling constant.
  • The only invariant and decomposable divergences for dimension D > 1 are f-divergences.
  • Statistical f-divergences are invariant under one-to-one or sufficient-statistic transformations of the sample space.
  • Invariant standard f-divergences induce the Fisher information matrix infinitesimally and yield the expected α-connections.
  • The expected α-connections are invariant connections induced by invariant statistical divergences, although their curvature depends on α and the statistical model.

3.11 Fisher-Rao expected Riemannian manifolds: (P, Pg)

The Fisher information matrix supplies the Riemannian metric for parametric distribution families, producing Fisher-Rao distances and geometry that is independent of α-representation.

  • The Fisher information matrix is used as the Riemannian metric tensor for regular parametric families of probability distributions.
  • The Fisher-Rao distance is the geodesic metric distance of the Fisher-Riemannian manifold.
  • Categorical, bivariate location-scale, and location families yield spherical, hyperbolic, and Euclidean Fisher-Riemannian manifolds, respectively.
  • The Riemannian structure of a parametric distribution family coincides with the self-dual conjugate-connection manifold induced by a symmetric f-divergence such as squared Hellinger divergence.
  • The α-representation rewrites densities and the Fisher information matrix through α-likelihood functions and α-score tangent bases.
  • The Fisher information matrix is α-independent despite the family of α-representations.

3.13 Dually flat spaces and canonical Bregman divergences

Dually flat spaces arise from strictly convex potentials and their Legendre-Fenchel conjugates, with canonical Bregman divergences recovering statistical distances such as forward and reverse KL.

  • A strictly convex C3 generator F defines a Hessian metric, while its convex conjugate F* defines the dual Hessian structure and coupled dual connections.
  • Canonical Legendre-Fenchel divergences associated with exponential-family log-normalizers recover reverse Kullback-Leibler divergence.
  • For exponential families, DKL[pθ1 : pθ2] equals the Bregman divergence with reversed parameter order, BF(θ2 : θ1).
  • For mixture families, the Bregman divergence generated by Shannon negentropy corresponds to the forward KL divergence.
  • The mixture-family entropy generally lacks a closed form because of the log-sum term, except when component distributions have pairwise disjoint supports.
  • Dually flat spaces can also be generated from homogeneous cones through the logarithm of their characteristic function.

4 Some applications of information geometry

Information geometry has broad applications across statistics, machine learning, signal processing, mathematical programming, and game theory, including natural gradient descent.

  • In statistics, information geometry is applied to asymptotic inference, expectation-maximization, and ARMA time-series models.
  • In machine learning, applications include restricted Boltzmann machines, neuromanifolds, and natural gradient methods.
  • Signal-processing applications include principal component analysis, independent component analysis, and non-negative matrix factorization.
  • Information geometry is also used for barrier functions in interior-point mathematical programming methods.
  • In game theory, information geometry is applied to score functions.
  • Natural gradient descent is presented as a celebrated application of information geometry.

Conjugate Connection Manifolds(M, g, ∇, ∇∗)

This section develops natural-gradient optimization on information manifolds, emphasizing coordinate invariance, computational approximations, and connections to dually flat geometry and mirror descent.

  • Natural gradient and Riemannian geometry: Natural gradient addresses parameterization dependence by selecting the steepest direction with respect to a Riemannian metric.It is invariant under invertible smooth coordinate transformations.
  • Natural gradient and Riemannian geometry: The Riemannian exponential map is often computationally intractable, so natural-gradient descent uses a Euclidean retraction as a first-order approximation.The retraction R_p(v)=p+v recovers the natural-gradient update rule.
  • Natural gradient and Riemannian geometry: Natural-gradient updates may leave the parameter manifold when the parameter domain is not all of R^D.Coordinate invariance of the gradient does not guarantee that every iterated location remains in Θ.
  • Dually flat spaces and mirror descent: Bregman mirror descent on a Hessian manifold is equivalent to natural-gradient descent on the dual Hessian manifold.The dual coordinates satisfy η=∇F(θ) and θ=∇F*(η).
  • Dually flat spaces and mirror descent: On dually flat spaces, natural gradient equals ordinary gradient descent after transforming to dual coordinates induced by a convex potential.The relation is NG∇_θL_θ=∇_ηL_η with η=∇_θF(θ).
  • Applications: The survey applies dually flat information manifolds to Bayesian hypothesis testing and clustering statistical mixtures.Chernoff information becomes a Bregman divergence for exponential families, while mixture-family geometry supports clustering with shared component distributions.

5 Conclusion: Summary, historical background, and perspectives

The survey concludes by organizing information geometry around dualistic structures, fundamental theorems, historical foundations, and divergence-based perspectives. It emphasizes that divergences can generate information manifolds and motivate broad applications beyond statistical models.

  • Summary: The dualistic structure couples a metric tensor with conjugate connections whose parallel transports preserve the metric.
  • Summary: The construction extends any conjugate-connection pair into a one-parameter family indexed by α.The pipeline is (M, g, ∇, ∇*) ⇒ (M, g, C) ⇒ (M, g, αC) ⇒ (M, g, ∇−α, ∇α).
  • Summary: The fundamental theorem characterizes dual constant-curvature manifolds, while dually flat manifolds provide dual potentials and global affine coordinates linked by Legendre-Fenchel transformation.
  • Summary: Information geometry extends beyond Fisher-Rao Riemannian modeling because information manifolds can arise from arbitrary divergences and need not be statistical.For symmetric divergences, induced conjugate connections coincide with Levi-Civita connection, but Fisher-Rao distance need not equal the divergence.
  • Perspectives: Metric distances and divergences differ structurally: a Riemannian distance is not generally a divergence, whereas its square is always a divergence.
  • Perspectives: The transformation T(u) = 1/(1+u) converts an unbounded metric into a bounded metric, whereas squaring a metric need not preserve metricity.
  • Historical background: Historically, Hotelling modeled parametric probability families as Riemannian manifolds using the Fisher Information Matrix as metric tensor, followed by Chentsov’s and Amari’s developments.
  • Perspectives: The survey highlights applications including Bayesian hypothesis testing and statistical mixture clustering, and recommends further study of generic divergence classes.

A Monte Carlo estimations of f-divergences

This appendix develops f-divergence estimation with Monte Carlo sampling and introduces extended f-divergences to address normalization-related problems. It also connects the resulting constructions to Bregman and Itakura-Saito divergences and to the Fisher information matrix.

  • Monte Carlo estimation: The Kullback-Leibler divergence is difficult to calculate in closed form for statistical mixtures, motivating Monte Carlo estimation with a proposal distribution.
  • Monte Carlo estimation: Monte Carlo estimators are consistent under mild conditions, with lim n→∞ cKL_n(p : q) = KL(p : q).
  • Monte Carlo estimation: The direct estimator can become negative because empirical samples may not satisfy the normalization condition, potentially disrupting algorithms that require nonnegative divergences.
  • Extended f-divergences: Extended f-divergences are introduced as a way to circumvent the potential negativity problem of direct Monte Carlo estimators.
  • Extended f-divergences: The extended Kullback-Leibler estimator can be interpreted as a sum of scalar Itakura-Saito divergences because the scalar divergence is scale-invariant.
  • Extended f-divergences: The extended f-divergence can be represented as an f-divergence for a transformed generator and interpreted through scalar Bregman divergences.
  • Fisher information connection: With normalization f′′(1) = 1, the standard f-divergence locally recovers the Fisher information quadratic form.

B The multivariate Gaussian family: An exponential family

The appendix presents the multivariate Gaussian family as an exponential family with ordinary, natural, and moment coordinates. It then relates Gaussian Kullback-Leibler divergence to Bregman divergences and, for equal covariances, to Mahalanobis distance.

  • Gaussian exponential family: The multivariate Gaussian family consists of distributions N(µ, Σ) with µ ∈ R^d and positive-definite covariance Σ.The family is also called the multivariate normal or MVN family.
  • Gaussian exponential family: The MVN density is written in canonical exponential-family form using compound vector-matrix natural parameters and sufficient statistics.
  • Dual coordinates: The log-normalizer is strictly convex and continuously differentiable, and its Legendre conjugate supplies the dual potential.
  • Dual coordinates: The appendix gives conversion formulas among ordinary parameters λ, natural parameters θ, and moment parameters η.
  • Gaussian divergences: The Gaussian Kullback-Leibler divergence has a closed form in terms of mean differences and covariance matrices.
  • Gaussian divergences: For equal covariance matrices, Gaussian Kullback-Leibler divergence equals half the squared Mahalanobis distance with precision matrix Σ^-1.
  • Gaussian divergences: Between members of the same exponential family, Kullback-Leibler divergence equals a Bregman divergence and its dual representation.

Notations

The notation appendix consolidates symbols for manifolds, tangent spaces, connections, divergences, potentials, coordinates, tensors, and statistical families used throughout the survey.

  • Divergences and potentials: Bregman and canonical divergences are defined through a strictly convex potential and its Legendre-Fenchel dual, while Chernoff information is separately specified.
  • Geometric objects: The notation distinguishes manifolds, submanifolds, tangent planes, tangent bundles, smooth functions, vector fields, and coordinate bases.
  • Differential geometry: Connections, geodesics, parallel transport, Christoffel symbols, curvature, and Lie brackets are listed as core differential-geometric notation.
  • Information geometry: The conjugate-connection framework uses the metric g, dual connections ∇ and ∇*, and the Amari-Chentsov cubic tensor C.
  • Information geometry: The notation identifies exponential families, mixture families, probability simplices, Fisher information matrices, expected α-connections, and statistical divergences.
Loading 1808.08271v2…