Source-linked AI summary
Adaptive Bayesian estimation using a Gaussian random field with inverse Gamma bandwidth
A. W. van der Vaart, J. H. van Zanten
TL;DR
The paper addresses adaptive nonparametric Bayesian estimation when the regularity of a multidimensional data-generating function is unknown. It rescales a smooth Gaussian random field with a Gamma variable, interpreting its reciprocal as a bandwidth, and studies the resulting posterior in density estimation, regression, and classification. The posterior contracts at the minimax rate up to a logarithmic factor across regularity levels, while the analysis assumes compactly supported parameters.
Problem
Nonparametric estimation quality depends on unknown function regularity, motivating Bayesian adaptation through a prior over bandwidth choices.
Method
The procedure uses a fixed prior formed by rescaling a smooth Gaussian random field with an independent Gamma-derived variable, whose reciprocal acts as a bandwidth.
Results
The posterior contracts at the minimax rate up to a logarithmic factor for every α > 0 in density estimation, and also adapts to specified infinitely smooth function classes.
Takeaways & Limitations
A single inverse Gamma bandwidth provides simultaneous adaptation across regularity levels without making the prior depend on α.
Takeaways & Limitations
The analysis considers only compactly supported parameters; without tail restrictions, full-space posterior consistency or contraction rates are not established.
Abstract
from arXiv · showhide
We consider nonparametric Bayesian estimation inference using a rescaled smooth Gaussian field as a prior for a multidimensional function. The rescaling is achieved using a Gamma variable and the procedure can be viewed as choosing an inverse Gamma bandwidth. The procedure is studied from a frequentist perspective in three statistical settings involving replicated observations (density estimation, regression and classification). We prove that the resulting posterior distribution shrinks to the distribution that generates the data at a speed which is minimax-optimal up to a logarithmic factor, whatever the regularity level of the data-generating distribution. Thus the hierachical Bayesian procedure, with a fixed prior, is shown to be fully adaptive.
1. Introduction.
The paper develops a fixed-prior Bayesian scheme that rescales a smooth Gaussian random field using an inverse Gamma bandwidth. It establishes adaptation across regularity levels and extends the framework to multidimensional density estimation, regression, and classification, while restricting parameters to compact support.
- Adaptive estimation traditionally uses bandwidth-indexed estimators, with Bayesian procedures instead placing a prior on the bandwidth and selecting it through the posterior.
- The proposed fixed prior rescales a smooth Gaussian random field, with the squared exponential process and inverse Gamma bandwidth as one possible choice.
- Rescaling by A makes the inverse 1/A act as a bandwidth, with larger A producing more randomness and accommodating less regular functions.
- A single inverse Gamma bandwidth is shown to adapt simultaneously to every regularity level α, including multivariate and certain infinitely smooth functions.
- The framework covers density estimation, fixed-design regression, and classification, and its abstract result also applies to broader Gaussian fields and bandwidth distributions.
- The analysis restricts parameters to compactly supported functions because posterior consistency on the full Euclidean space requires tail restrictions.
2. Main results.
The paper applies a fixed, randomly rescaled Gaussian-process prior to density estimation, regression, and classification. Across these settings, posterior contraction adapts to unknown regularity at minimax rates up to logarithmic factors.
- 2. Main results.: The main results cover i.i.d. density estimation, fixed-design regression, and binary classification.The three settings use transformed or direct rescaled Gaussian-process priors as appropriate.
- 2.1. Density estimation.: The density prior is obtained by exponentiating and renormalizing a randomly rescaled Gaussian process.The posterior is evaluated under the frequentist assumption that the observations are sampled from the true density f0.
- 2.1. Density estimation.: For α-Hölder log-densities, the posterior contracts at n^-α/(2α+d) multiplied by a logarithmic factor.This is the minimax rate n^-α/(2α+d) up to the stated logarithmic factor.
- 2.1. Density estimation.: For infinitely smooth log-densities, the contraction rate is n^-1/2(log n)^(d+1) when r ≥ 2 and n^-1/2(log n)^(d+1+d/(2r)) when r < 2.The rate improves as r increases but does not improve beyond r = 2 for the squared exponential process.
- 2.2. Regression and 2.3. Classification.: The same contraction assertions as in density estimation hold for regression when w0 = f0 and for classification when w0 = Ψ^-1(r0).Regression uses the rescaled Gaussian process directly, while classification applies a logistic or normal link function.
3. Rescaled Gaussian fields.
The theoretical analysis studies a Gaussian field rescaled by an independent positive random variable and establishes the small-ball and support-complexity conditions needed for posterior contraction. Under spectral and scaling assumptions, the resulting rates apply across all three statistical settings.
- 3. Rescaled Gaussian fields.: The field is assumed to have a spectral measure with subexponential tails and a Lebesgue density satisfying a radial monotonicity condition.The squared exponential process belongs to this class, with spectral density proportional to exp(-||λ||^2/4).
- 3. Rescaled Gaussian fields.: The rescaled process is WA(t) = WAt on [0,1]^d, viewed as a random element of C[0,1]^d with the uniform norm.A is positive and stochastically independent of the centered homogeneous Gaussian field W.
- 3. Rescaled Gaussian fields.: A Gamma distribution for Ad satisfies the scaling-density assumption with q = 0.The random scaling variable is used to provide the inverse-Gamma bandwidth mechanism.
- 3. Rescaled Gaussian fields.: For α-smooth functions, the abstract theorem gives εn = n^-α/(2α+d)(log n)^κ1 and ε̄n = Kεn(log n)^κ2 under its stated conditions.The exponents κ1 and κ2 depend on d, α, and q.
- 3. Rescaled Gaussian fields.: For analytic-type functions, the theorem gives εn = Kn^-1/2(log n)^(d+1)/2 when r ≥ ν and an additional d/(2r) logarithmic exponent when r < ν.These bounds require a lower exponential bound on the spectral density.
- 3. Rescaled Gaussian fields.: The abstract conditions map one-to-one to general posterior contraction conditions in density estimation, regression, and classification.This yields contraction at rate εn ∨ ε̄n for each setting; q = d + 1 gives a slightly better logarithmic power, while q = 0 corresponds to a Gamma prior.
4. Auxiliary results.
The auxiliary results characterize the rescaled Gaussian process through its RKHS, approximation properties, entropy bounds, and concentration probabilities. These tools connect smoothness of the target function to prior behavior under rescaling.
- Concentration analysis: For fixed a, approximation and small-ball probabilities for W_a are controlled through the RKHS geometry and metric entropy.The concentration function combines centered small probabilities with the RKHS approximation distance to w0.
- Density argument: Analytic continuation shows that polynomials are dense in L2(µ) under the stated spectral conditions, completing the relevant RKHS characterization.Orthogonality to all polynomials implies vanishing of the associated Fourier transform and hence the function almost everywhere.
- RKHS characterization: The RKHS of the rescaled process W_a is obtained from the original process through the scaling map h ↦ (t ↦ h(at)).This map is an isometry between the RKHS on [0,a]^d and the RKHS of W_a.
- Supersmooth targets: If the spectral density is bounded below by exp(−D3∥λ∥^ν), analytic targets with regularity r ≥ ν belong to H_a for sufficiently large a with uniformly bounded norm.For r < ν, the corresponding RKHS approximation requires bounds depending on a.
- Entropy bounds: The RKHS unit ball admits finite ε-nets constructed from piecewise polynomials, yielding entropy bounds used in the concentration analysis.The covering construction partitions the domain and approximates functions locally by polynomials.
5. Proof of Theorem 3.1.
The proof selects bounds and tuning sequences for the rescaled Gaussian prior according to the target’s smoothness class. These choices control prior concentration and sieve probabilities at the required posterior contraction rates.
- Hölder smoothness: For Hölder targets, choosing ε_n as a multiple of n^−α/(2α+d) times a logarithmic factor satisfies the required concentration conditions.The proof uses ε > C a^−α and bounds the concentration function by a multiple of nε_n^2.
- Hölder smoothness: The Hölder proof chooses r_n minimally and then M_n so that tail and sieve conditions hold, with the remaining condition automatic for large n.The resulting entropy contribution is bounded using powers of r_n and log n.
- Analytic smoothness: For analytic targets with r ≥ ν, the proof combines RKHS approximation and prior-tail bounds to obtain exponentially small complement probabilities.The choices of B, r, and M ensure P(W_A ∉ B) is bounded by a multiple of exp(−C0 n ε_n^2).
- Analytic smoothness: When r < ν, the contraction scale includes an additional logarithmic exponent depending on d and r.The proof obtains ε_n as a large multiple of n^−1/2(log n)^{d/(2r)+(d+1)/2}.
- Entropy control: The entropy calculations bound the sieve complexity by terms involving r_n^d and logarithmic factors, which are then matched to n ε_n^2.These bounds are verified after selecting r, M, and ε according to the target regularity regime.