Source-linked AI summary
Adaptive Bayesian multivariate density estimation with Dirichlet mixtures
Weining Shen, Surya T. Tokdar, Subhashis Ghosal
TL;DR
The paper addresses technical gaps in rate-adaptive nonparametric Bayesian density estimation. It establishes adaptive convergence results for multivariate normal-mixture priors, including minimax-optimal rates across smoothness classes.
Problem
Rate-adaptation results for nonparametric Bayesian density estimation faced two major technical difficulties involving concentration rates for mixture priors.
Method
The paper studies Dirichlet process location mixtures of normals and establishes rate-adaptation properties for multivariate density estimation.
Results
The resulting convergence rate is n^-β/(2β+d)(log n)^t, and without the logarithmic factor it is minimax optimal for the β-Hölder class; anisotropic rates are also minimax optimal.
Takeaways & Limitations
Normal-kernel mixtures achieve optimal convergence rates universally across smoothness levels.
Takeaways & Limitations
Theorem 1 gives the optimal rate only for d ≥2, and κ has a substantial impact on rates.
Abstract
from arXiv · showhide
We show that rate-adaptive multivariate density estimation can be performed using Bayesian methods based on Dirichlet mixtures of normal kernels with a prior distribution on the kernel's covariance matrix parameter. We derive sufficient conditions on the prior specification that guarantee convergence to a true density at a rate that is optimal minimax for the smoothness class to which the true density belongs. No prior knowledge of smoothness is assumed. The sufficient conditions are shown to hold for the Dirichlet location mixture of normals prior with a Gaussian base measure and an inverse-Wishart prior on the covariance matrix parameter. Locally Hölder smoothness classes and their anisotropic extensions are considered. Our study involves several technical novelties, including sharp approximation of finitely differentiable multivariate densities by normal mixtures and a new sieve on the space of such densities.
1. INTRODUCTION
The paper addresses missing multivariate rate-adaptation results for Dirichlet process mixture density estimators and establishes optimal convergence across Hölder smoothness classes without prior smoothness knowledge.
- 1. INTRODUCTION: Adaptive rate results for multivariate Dirichlet process mixture density estimators had remained unavailable beyond univariate density estimation.The main technical obstacles involve adaptive prior concentration for nonnegative mixture densities and scalable, adaptive sieves.
- 1. INTRODUCTION: The study contributes rate-adaptation theory for kernel-based multivariate density estimation methods.The authors state that these results are new for any kernel-based multivariate density estimation method.
- 1. INTRODUCTION: The method uses normal kernels with locations drawn from a Dirichlet process with Gaussian base measure and an inverse-Wishart prior on the common covariance matrix.Rate adaptation is established for Hölder smoothness classes under this commonly used specification.
- 1. INTRODUCTION: For locally β-Hölder densities in d dimensions, the posterior rate is n^-β/(2β+d)(log n)^t, minimax optimal without the logarithmic factor.The exponent t depends on β, d, and tail properties of the true density.
- 1. INTRODUCTION: For anisotropic Hölder densities, the rate is n^-β0/(2β0+d) times a log n factor, with β0 the harmonic mean of the directional smoothness coefficients, and is minimax optimal.The result covers densities with different smoothness orders along the coordinate axes.
- 1. INTRODUCTION: A single Dirichlet process mixture of normal kernels achieves optimal convergence rates universally across smoothness levels.This contrasts with kernel estimators whose optimal bandwidth choices require knowing the smoothness level.
2. POSTERIOR CONVERGENCE RATES FOR DIRICHLET MIXTURES
The paper establishes posterior convergence rates for Dirichlet mixtures of multivariate normals under locally Hölder assumptions, with rates governed by smoothness, dimension, and covariance-prior tail behavior.
- Prior specification: The prior is a Dirichlet process location mixture of normals with Gaussian-supported locations and a covariance distribution satisfying tail and local-mass conditions.The covariance prior must control both large and small eigenvalues and place sufficient mass near suitable covariance matrices.
- Prior specification: An inverse-Wishart covariance prior satisfies the required covariance conditions, including the full-support specification allowed by the theory.The inverse-Wishart result gives κ=2.
- Posterior convergence rates: The posterior contracts in Hellinger or L1 distance at n^-β/(2β+d*) times a logarithmic factor, where d*=max(d,κ).Here β is the Hölder smoothness and κ is determined by the covariance prior.
- Proof strategy: The proof verifies general posterior-contraction conditions using prior thickness, entropy bounds, and a sieve of probability densities.The sieve controls the model class while Kullback–Leibler prior mass supplies thickness around the true density.
- Posterior convergence rates: When κ=1, the rate is minimax-optimal up to a logarithmic factor for the β-Hölder class.The benchmark minimax rate is n^-β/(2β+d).
- Prior-dependent scope: For the standard inverse-Wishart specification with κ=2, the theorem yields the optimal rate only when d≥2.The covariance-prior parameter κ also affects convergence rates for anisotropic densities.
- Finite mixtures: The same convergence-rate framework applies to finite normal-mixture priors, with an additional logarithmic-rate penalty depending on τ.The finite-mixture rate uses d*=max(d,κ) and requires t greater than the stated threshold.
3. PRIOR THICKNESS RESULTS
The paper develops sharp normal-mixture approximations for locally Hölder densities and uses them to establish prior-thickness results under Hellinger distance. Compactly supported density approximations and matrix-valued covariance handling connect these constructions to posterior-rate analysis.
- Approximation: Functions in Cβ,L,τ0 can be approximated by normal mixtures with pointwise error of order σβ.Lemma 2 bounds the approximation error by MβL(x)σβ.
- Approximation: For probability densities, a density hσ is constructed from the transform Tβ,σf so that Kσhσ achieves σβ-order approximation.The transform itself need not be a density and may be negative, so the construction replaces it with a valid mixing density.
- Novelty: The approximation proof extends earlier results by using multivariate Taylor expansions on f0 and establishing bounds under Hellinger distance.The authors identify this as a main technical difference from the related univariate work.
- Prior thickness: Under tail and smoothness conditions, a compactly supported density h̃σ satisfies dH(f0, Kσh̃σ) ≤ K0σβ.Its support lies in a ball of radius a0{log(1/σ)}^τ, enabling subsequent prior-thickness calculations.
- Prior thickness: The subsequent approximation argument handles non-compactly supported f0 and matrix-valued covariance parameters.These calculations support the final prior-thickness theorem for the normal-mixture model.
0. Fix
The construction approximates the compactly supported target by discretizing its mixing distribution and controlling perturbations from discretization and covariance changes. These controls yield Hellinger bounds for the resulting normal mixtures.
- Partition construction: The mixing distribution is discretized over separated small balls and extended to a partition of the support region and all of Rd.The construction uses centers separated by σ ε̃_n^(2b1) and partitions with controlled diameters.
- Mixing-measure control: The selected mixing measures assign controlled masses to partition cells, including a lower bound on the minimum cell mass.These constraints define the class Pσ used in the approximation argument.
- Covariance control: The covariance class Sσ restricts eigenvalues of Σ^-1 to a narrow interval around σ^-2.This restriction controls determinant and quadratic-form differences between pF,Σ and the isotropic mixture.
- Approximation construction: The discretized mixture and covariance perturbations jointly yield dH(f0, pF,Σ) bounded by a constant multiple of σβ.For F in the selected class and Σ in Sσ, the proof combines approximation, discretization, and covariance bounds.
- Tail control: Tail conditions control the contribution outside the compact support and supply finite moments needed in the Hellinger analysis.The proof applies the tail condition to bound the probability outside the support region.
4. SIEVE CONSTRUCTION
The paper introduces an explicit sieve for Dirichlet-mixture densities and bounds both its entropy and the prior probability of its complement. This sieve is the main tool for adaptive posterior convergence rates.
- Construction: The sieve is built from a stick-breaking representation of the Dirichlet process.Its complement probability is bounded using the resulting finite-dimensional restrictions and tail controls.
- Sieve role: The explicit sieve provides entropy and prior-complement bounds needed to obtain adaptive posterior convergence rates.The authors identify this result as the main tool for the adaptive convergence analysis.
- Rate calibration: The sieve supports rates εn = n^-γ(log n)^t for 0 < γ ≤ 1/2, with γ = β/(2β + d*) for β-Hölder classes.The parameter choice connects the general sieve rate to the Hölder-class convergence rate.
- Entropy control: For the constructed sieve, the entropy bound is controlled by terms of order n^(1−2γ)(log n)^(2t) together with logarithmic terms.The theorem establishes the required entropy condition for sufficiently large n.
5. ANISOTROPIC H ¨OLDER FUNCTIONS
The anisotropic extension allows different smoothness orders across coordinate axes and yields rates based on the effective anisotropic dimension. It improves over isotropic analysis when the density is genuinely anisotropic, subject to covariance-prior limitations.
- Class definition: Anisotropic classes assign axis-specific smoothness βj = β/αj and include mixed derivatives through orders determined by each βj.The isotropic class is recovered when α = (1,...,1).
- Comparison with isotropic rates: When anisotropy is present, the anisotropic theorem is sharper because isotropic analysis uses the smaller smoothness index β/αmax.The isotropic index is strictly smaller than β unless all αj equal 1.
- Anisotropic rate: The anisotropic posterior rate is εn = n^-β/(2β+d*) (log n)^t in Hellinger or L1 distance.The exponent uses d* and the logarithmic power t specified by the theorem.
- Method: The anisotropic construction replaces the single bandwidth σ with bandwidth σαj along axis j.This modification is used to approximate densities with coordinate-dependent smoothness.
- Prior dependence: With a standard inverse-Wishart prior, optimal rates are recovered only for limited anisotropy, whereas diagonal inverse-gamma covariance priors support optimal rates for any dimension and anisotropy.The standard prior has κ = 2; the diagonal alternative has κ ≤ 1.
APPENDIX A. PROOFS
The appendix proves approximation, metric, prior-mass, and sieve results underlying the Bayesian multivariate density-estimation theory, including anisotropic extensions.
- Prior concentration: The inverse-Wishart analysis controls eigenvalue probabilities and extends the identity-scale calculation to general positive-definite scale matrices.The proof uses chi-square tail bounds and transformed inverse-Wishart distributions.
- Approximation: The appendix constructs truncated approximating densities whose Hellinger distance from the target is bounded by Kσ^2β.The construction also verifies normalization, positivity on a high-density region, and integrability under derivative conditions.
- Sieve construction: The covariance-discretization proof bounds the L1 error from eigenvalue and eigenvector approximation, with the first term bounded by 2ε.The argument decomposes mixture error into covariance and mixing-measure components before controlling each term.
- Sieve construction: A finite 6ε-net for normal mixtures is obtained by discretizing locations, weights, eigenvalues, and orthogonal matrices, with cardinality bounded by a product of covering factors.The location, simplex, and orthogonal-matrix nets have cardinalities controlled by {a/(σ0ε)}^d, ε^-H, and δ^-d(d−1)/2, respectively.
- Prior concentration: The Dirichlet-process representation converts the prior into a discrete normal mixture through stick-breaking weights and independently sampled locations.The resulting prior-mass bounds combine location, weight, and covariance contributions.
APPENDIX B. SUPPLEMENTARY RESULTS
The supplementary results provide discrete normal-mixture approximations, metric inequalities, and regularity consequences needed for the main posterior-convergence arguments.
- Discrete approximation: A discrete mixing distribution with at most D[{(a/σ) ∨ 1} log(1/ε)]^d support points approximates a Gaussian-smoothed distribution in sup norm and L1.The bounds are ∥pP0,σ − pFσ,σ∥∞ ≲ ε/σ^d and ∥pP0,σ − pFσ,σ∥1 ≲ ε{log(1/ε)}^1/2.
- Auxiliary lemmas: The supplementary lemmas extend mixture approximation, moment matching, and metric inequalities to d dimensions.The moment-matching construction uses at most {(2k−2)^d+1} support points, and the resulting dimensional dependence propagates into Nσ,ε.
- Discrete approximation: The support points can be placed on a regular grid while preserving the approximation up to ε-scale sup-norm and L1 errors.Moving support points to the grid contributes at most a constant times ε^2/σ^d in sup norm and a constant times ε in L1.
- Regularity conditions: Regularity assumptions on log f0, polynomial moment conditions, and the tail condition imply that f0 belongs to a suitable locally Hölder class.The result constructs an envelope L and a parameter τ0 from the stated assumptions.