Source-linked AI summary

Convergence of latent mixing measures in finite and infinite mixture models

XuanLong Nguyen

arXiv:1109.3250v5math.STstat.ML

TL;DR

The paper addresses how latent mixing measures converge, beyond the more commonly studied convergence of mixture densities. It uses Wasserstein metrics, relates them to f-divergences through identifiability conditions, and establishes posterior contraction rates for finite and Dirichlet process mixtures. These metrics also connect convergence of discrete mixing measures to convergence of their atoms.

  • Problem

    Convergence results for latent mixing measures are rare, while prior work was mainly limited to univariate finite mixtures with a known component bound.

  • Method

    The paper analyzes Wasserstein distances on mixing measures, their relationships with divergences between mixture densities, and posterior contraction under separate prior and likelihood conditions.

  • Results

    Posterior contraction rates are established for finite multivariate mixtures and infinite Dirichlet process mixtures, including (log n)1/4n−1/4 for bounded-support finite mixtures in Rd.

  • Takeaways & Limitations

    For discrete measures, Wasserstein convergence implies convergence of supporting atoms, giving a clustering interpretation in terms of converging typical behaviors.

  • Takeaways & Limitations

    The infinite-mixture identifiability theorem is restricted to convolution mixture models, although bounded-support assumptions can be replaced by bounded moments with a weaker ordinary-smooth bound.

Abstract

from arXiv · show

This paper studies convergence behavior of latent mixing measures that arise in finite and infinite mixture models, using transportation distances (i.e., Wasserstein metrics). The relationship between Wasserstein distances on the space of mixing measures and f-divergence functionals such as Hellinger and Kullback-Leibler distances on the space of mixture distributions is investigated in detail using various identifiability conditions. Convergence in Wasserstein metrics for discrete measures implies convergence of individual atoms that provide support for the measures, thereby providing a natural interpretation of convergence of clusters in clustering applications where mixture models are typically employed. Convergence rates of posterior distributions for latent mixing measures are established, for both finite mixtures of multivariate distributions and infinite mixtures based on the Dirichlet process.

1. Introduction.

The paper develops Wasserstein-based analysis for convergence of latent mixing measures in finite and infinite mixture models, relating these distances to divergences between mixture densities. It establishes identifiability results and posterior contraction rates, including finite multivariate mixtures and Dirichlet process mixtures.

  • Motivation: Mixing measures combine simple models into richer statistical models and can be random and infinite-dimensional under Bayesian nonparametric priors.These constructions support modeling complex and high-dimensional data.
  • Motivation: Existing Bayesian mixture analyses primarily study posterior convergence of data densities, while convergence results for latent mixing measures are comparatively rare.Earlier mixing-measure results were limited to univariate finite mixtures with a known bound on the number of components.
  • Wasserstein framework: Wasserstein distances provide a natural metric because discrete-measure comparisons can be formulated as minimum matching between their supporting atoms.They inherit the metric structure of the atom space and support interpretation of atom convergence as convergence of heterogeneous behaviors.
  • Identifiability: The paper relates Wasserstein distances between mixing measures to total variation, Hellinger, and Kullback–Leibler divergences between mixture densities under identifiability conditions.It gives an upper bound from Wasserstein distances to f-divergences and bounds W2 in terms of divergences for bounded-atom and convolution mixture settings.
  • Posterior convergence: Posterior distributions of latent mixing measures are shown to contract under W2 using separate prior and likelihood conditions, typically stronger than weak convergence from Hellinger metrics.The results are established through the standard posterior-contraction approach of Ghosal, Ghosh, and van der Vaart.
  • Posterior convergence: For bounded-support finite mixtures in Rd, the posterior rate is (log n)1/4n−1/4, while Dirichlet process mixtures achieve smoothness-dependent rates for ordinary- and supersmooth likelihoods.The finite-mixture rate is minimax-optimal up to a logarithmic factor; ordinary smoothness yields (log n/n)γ for any γ < 2/((d+2)(4+(2β+1)d)), and supersmoothness yields (log n)−1/β.

2. Transportation distances for mixing measures.

The paper defines transportation distances for discrete mixing measures and relates them to divergences between mixture distributions. Under identifiability and regularity conditions, Wasserstein convergence controls atom convergence and can be quantitatively related to mixture-density distances.

  • Definitions: A discrete mixing measure assigns probabilities to atoms in Θ, with finite or countably infinite support.Finite-support measures form G(Θ), while Ḡ(Θ) also includes countably infinite support.
  • Transportation distances: Wasserstein distance transports mass between atoms using a distance ρ on Θ, and W1 and W2 are standard cases when ρ is the metric on Θ.The transportation construction uses couplings satisfying the marginal constraints of the two proportion vectors.
  • Relationship to mixture divergences: Any finite f-divergence between mixture densities pG and pG′ is bounded above by the corresponding transportation distance between G and G′.This inequality also supports testing bounds and lower bounds for small Kullback–Leibler balls.
  • Identifiability and convergence: Identifiability conditions are necessary because mixture-density convergence alone need not imply transportation convergence of the latent mixing measures.The paper distinguishes finite identifiability from the stronger conditions used for convergence-rate results.
  • Finite mixtures: Strong identifiability with differentiability and regularity conditions yields a positive separation between mixture-density divergence and nearby W2 mixing measures around a finite-atom target.For a fixed G0 in Gk(Θ), the theorem establishes the displayed positive lower bound as ε approaches zero.

3. Convergence of posterior distributions of mixing measures.

The paper develops a Wasserstein-based framework for posterior convergence of latent mixing measures, connecting Wasserstein separation to mixture-density divergences through identifiability and Hellinger information. General contraction theorems use Wasserstein entropy, prior concentration, Kullback–Leibler support, and likelihood conditions.

  • The posterior-contraction framework establishes consistency in W2 neighborhoods of the true mixing measure and derives convergence rates.The general results use tests, entropy bounds, Kullback–Leibler support, and Hellinger information.
  • Wasserstein-based conditions are stated directly for mixing measures and can be separated into prior and likelihood requirements that are simpler to verify.This contrasts with conditions formulated through f-divergences on mixture densities.
  • Finite identifiability ensures positive Hellinger information away from the truth, while strong identifiability yields a lower bound proportional to r^4 for finite-support classes.These properties connect Wasserstein separation of mixing measures to distinguishability of their mixture distributions.
  • For ordinary smooth likelihoods, the Hellinger information grows polynomially in r, whereas supersmooth likelihoods yield an exponential lower bound.The corresponding forms are Ψ̄G(Θ)(r) ≥ c(d,β)r^(4+(2β+1)d′) and Ψ̄G(Θ)(r) ≥ exp[−c(d,β)r^−β].
  • Wasserstein entropy for finite- and infinite-support measure classes is bounded using covering numbers of the parameter space.The bounds differ according to whether the class has at most k support points or unrestricted discrete support.

4. Examples.

The examples section applies the general posterior-contraction theory to finite mixtures with known finite support size, under compactness and strong identifiability assumptions.

  • The paper presents these finite-mixture results alongside infinite mixtures based on the Dirichlet process as two principal application classes.The supplied examples-section passages specifically introduce the finite-mixture setting.
  • The finite-mixture example considers a prior on discrete measures in G_k(Θ), with finite known k and Θ a subset of R^d.The true mixing measure is assumed to belong to the same finite-support class.
  • The example requires Θ to be compact and the likelihood family f(·|θ) to be strongly identifiable.These are the stated assumptions used to obtain the convergence rate.

Assumptions A.

Under assumptions A, the finite-mixture posterior contracts in W2 at a near-minimax rate, with identifiability, regular likelihood behavior, and separated nonzero component weights supporting the result.

  • Assumption (A4) requires all mixing weights to be bounded away from zero and pairwise atom distances to be bounded away from zero.These conditions prevent vanishing components and coincident support points within the finite-mixture class.
  • Gaussian likelihoods with mean parameter θ satisfy assumptions (A1) and (A2), while a suitably regular prior can satisfy (A3).The paper notes that a prior behaving like a uniform distribution up to a multiplicative constant is sufficient for (A3).
  • Strong identifiability and assumption (A2) provide the lower bound Ψ_Gn(ε) ≥ cε^4 used in the contraction proof.This links Wasserstein separation to the likelihood-based testing argument.
  • (log n)^1/4 n^−1/4 is the finite-mixture posterior contraction rate in the stated argument, up to a logarithmic factor from the minimax rate n^−1/4.The comparison is stated for univariate finite mixtures.

Assumptions B.

Under assumptions B, a Dirichlet-process prior yields W2 posterior contraction for infinite mixtures, with rates determined by parameter dimension and whether the likelihood is ordinary smooth or supersmooth.

  • Theorem 6 establishes a sequence β_n → 0 such that posterior mass outside the W2 ball of radius β_n around G0 vanishes in PG0-probability.The result assumes (B1), (B2), and the likelihood smoothness conditions from Theorem 2.
  • For ordinary smooth likelihoods, the proof starts with ε_n as a large multiple of (log n/n)^(1/(d+2)) and sets β_n = M_n ε_n.The construction verifies entropy, prior concentration, and testing conditions before applying the general theorem.
  • Under ordinary smoothness, R_Ḡ(Θ)(t) = t^(1/(4+(2β+1)d+δ)), with δ an arbitrarily positive constant.This inverse Hellinger-information rate determines the contraction scaling.
  • Under supersmoothness, R_Ḡ(Θ)(t) = (1/log(1/t))^(1/β), yielding β_n ≍ (log n)^−1/β.The rate is logarithmic in n under the stated supersmooth likelihood condition.

5. Proofs.

The proofs establish Wasserstein identifiability through contradiction and strong identifiability, derive bounds relating mixing-measure Wasserstein distances to mixture-distribution divergences, and construct posterior contraction tests.

  • Wasserstein identifiability results: Theorem 1 uses subsequences, coefficient rescaling, and strong identifiability to rule out nonvanishing discrepancies between mixing measures.The argument handles coefficients that may diverge by normalizing them and derives a contradiction when the limiting linear combination cannot vanish almost everywhere.
  • Bounds between metrics: Mollification bounds W2(G,G′) by comparing convolved mixing measures and controlling the convolution error through kernel moments and coupling.The coupling gives W2(G,G∗Kδ)=O(δ), while Fourier arguments control the convolved mixture distributions.
  • Bounds between metrics: Under Fourier-tail assumptions on the mixture kernel, the proof obtains W2(G,G′) ≤ C(d,β,r)V(pG,pG′)^r for every r < 4/(4 + (2β + 1)d).The exponent approaches 4/(4 + (2β + 1)d) as the available moment order s increases.
  • Wasserstein identifiability results: Finite identifiability plus compactness converts vanishing Hellinger distance between mixture distributions into convergence in W2 for mixing measures.A separated sequence would have distinct compactness limits with identical mixture densities, contradicting finite identifiability.
  • Posterior contraction: The posterior contraction proofs construct tests separating W2 neighborhoods by Hellinger distance, using convexity or maximal Wasserstein packings.The testing bounds are based on the infimum Hellinger distance over separated mixture distributions and metric packing numbers.
Loading 1109.3250v5…