Source-linked AI summary
Frequentist Consistency of Variational Bayes
Yixin Wang, David M. Blei
TL;DR
Scalable posterior inference is needed because exact Bayesian computation can be intractable and MCMC convergence can be slow for large datasets, while rigorous VB theory remains limited. The paper connects VB to frequentist variational approximations and proves a variational Bernstein–von Mises theorem. It establishes asymptotic normality of the VB posterior and consistency and asymptotic normality of the VBE, with mean-field approximations retaining marginal behavior but missing posterior dependence.
Problem
Exact posterior computation can be intractable, MCMC convergence can be slow as data sizes grow, and few rigorous theoretical guarantees exist for VB.
Method
The paper connects the VB posterior to frequentist variational approximations through the variational log likelihood and analyzes this connection with a variational Bernstein–von Mises theorem.
Results
The VB posterior converges in total variation to the KL minimizer of a normal distribution centered at the truth, while the VBE is consistent and asymptotically normal.
Takeaways & Limitations
VB has theoretical soundness under the paper’s conditions, although mean-field Gaussian posteriors recover marginal variance but not off-diagonal covariance terms and are underdispersed.
Takeaways & Limitations
The theory assumes a global optimum, whereas VB optimization typically finds a local optimum; characterizing these local optima requires further study.
Abstract
from arXiv · showhide
A key challenge for modern Bayesian statistics is how to perform scalable inference of posterior distributions. To address this challenge, variational Bayes (VB) methods have emerged as a popular alternative to the classical Markov chain Monte Carlo (MCMC) methods. VB methods tend to be faster while achieving comparable predictive performance. However, there are few theoretical results around VB. In this paper, we establish frequentist consistency and asymptotic normality of VB methods. Specifically, we connect VB methods to point estimates based on variational approximations, called frequentist variational approximations, and we use the connection to prove a variational Bernstein-von Mises theorem. The theorem leverages the theoretical characterizations of frequentist variational approximations to understand asymptotic properties of VB. In summary, we prove that (1) the VB posterior converges to the Kullback-Leibler (KL) minimizer of a normal distribution, centered at the truth and (2) the corresponding variational expectation of the parameter is consistent and asymptotically normal. As applications of the theorem, we derive asymptotic properties of VB posteriors in Bayesian mixture models, Bayesian generalized linear mixed models, and Bayesian stochastic block models. We conduct a simulation study to illustrate these theoretical results.
1 Introduction
The paper develops a theoretical account of variational Bayes, connecting its posterior to variational approximations and proving frequentist consistency and asymptotic normality. It shows that the VB posterior approaches a normal-distribution KL minimizer centered at the truth, while the variational Bayes estimate is consistent and asymptotically normal.
- Motivation and setup: VB offers a scalable optimization-based alternative to MCMC for posterior inference when exact computation is intractable and data sizes are large.VB selects the distribution closest to the exact posterior under KL divergence, typically using the ELBO as the optimization objective.
- Motivation and setup: Mean-field VB factorizes the latent-variable distribution into n + d marginal distributions, separating local variables z from global parameters θ.The framework allows local variables to grow with sample size while the number of global variables remains fixed.
- Theoretical strategy: The paper connects the VB posterior to the variational log likelihood through the VB ideal and frequentist variational approximations.The VB ideal provides an intermediate object whose theoretical properties are established before analyzing the VB posterior.
- Main results: The variational Bernstein–von Mises results show that the VB posterior converges to the KL minimizer of a normal distribution centered at the truth, under consistency of the VFE.The VB ideal and its KL minimizer establish the intermediate asymptotic results used for the final posterior theorem.
- Main results: The variational Bayes estimate is consistent and asymptotically normal when the relevant frequentist variational estimate is consistent.For full-rank Gaussian families, the posterior recovers the true mean and covariance; for mean-field Gaussian families, it recovers the true mean and marginal variance but not off-diagonal terms.
- Scope and contribution: The paper broadens variational-method theory beyond treatments restricted to specific models and priors by proving a variational Bernstein–von Mises theorem.The analysis applies to broader model classes and is illustrated through Bayesian mixture, generalized linear mixed, and stochastic block models.
2 The VB ideal
The VB ideal is analyzed as a posterior under a frequentist variational model, enabling Bernstein–von Mises results for its consistency and asymptotic normality. Its KL minimizers and the VB posterior inherit convergence toward normal limits centered at the truth.
- The VB ideal is the classical posterior under the frequentist variational model, while the VFE is its classical maximum likelihood estimate.
- The analysis uses prior mass, consistent testability, and local asymptotic normality assumptions for the variational log likelihood.The normalizing sequence δn determines posterior convergence rates; in one dimension, δn commonly equals 1/√n.
- The VB ideal converges in total variation to a sequence of normal distributions.
- The rescaled VB ideal is asymptotically normal with a random mean and covariance determined by the variational model.The random mean is bounded in probability and reflects randomness in data drawn from the true data-generating measure.
- The KL minimizer of the VB ideal is consistent and converges in total variation to the KL minimizer of a normal distribution.
- The variational Bernstein–von Mises theorem transfers these consistency and asymptotic-normality properties from the KL minimizer to the VB posterior.
3 Frequentist consistency of variational Bayes
The paper connects the practical VB posterior to KL minimization and proves its frequentist consistency and asymptotic normality under a LAN condition. Results differ by variational family: full-rank Gaussian VB recovers the asymptotic normal distribution, while factorizable Gaussian VB is underdispersed.
- VB posterior and KL minimizers: The VB posterior is characterized by a profiled ELBO whose optimizer is asymptotically equivalent to the KL minimizer of the VB ideal.The equivalence holds under mild tail conditions on the variational family.
- Main theorem: The variational Bernstein–von Mises theorem establishes that the VB posterior is consistent and converges to the KL minimizer of a normal distribution.After centering at the true parameter and scaling by the convergence rate, the limiting posterior is normal for mean-field families.
- Main theorem: The variational expectation is consistent and asymptotically normal, converging in distribution to the mean of the KL minimizer.Its limiting behavior follows from the convergence of the VB posterior and the continuity of posterior means.
- Gaussian variational families: Under a full-rank Gaussian variational family, VB is consistent, asymptotically normal, and accurately recovers the LAN-implied asymptotic normal distribution.The full-rank family avoids the covariance restriction imposed by factorizable Gaussian approximations.
- Gaussian variational families: Under a factorizable Gaussian family, VB retains the correct mean but underestimates covariance and is therefore underdispersed.The factorizable KL minimizer has a diagonal covariance matching the precision and no greater entropy than the original Gaussian.
- LAN condition: The theory relies on Assumption 1.3, the LAN expansion of the variational log likelihood, whose derivation is model-specific when local latent variables are present.For regular parametric models without local latent variables, the variational log likelihood equals the ordinary log likelihood.
4 Applications
The paper applies its VB asymptotic theory to Gaussian mixture models, Poisson GLMMs, and stochastic block models, deriving consistency and asymptotic normality under model-specific conditions.
- Proof strategy: Across these applications, the theoretical analysis relies on consistent testability and local asymptotic normality of the variational log likelihood.The paper derives these ingredients using model-specific prior results and then invokes its central theorems.
- Applications: The authors apply their general VB argument to Bayesian mixture models, generalized linear mixed models, and stochastic block models.They leverage existing asymptotic results for frequentist variational approximations in each model class.
- Bayesian mixture models: Bayesian Gaussian mixture models use global component means and local observation-level cluster assignments, with inference targeting the posterior of the global means.The model specification is permutation invariant, so the result holds up to permutations of the K components.
- Bayesian generalized linear mixed models: The Bayesian Poisson GLMM models grouped counts using global intercept, slope, and variance parameters plus group-specific Gaussian random intercepts.The paper establishes asymptotic properties of VB for this model under regularity conditions, including m = O(n^2).
- Bayesian stochastic block models: For stochastic block models, latent class assignments generate Bernoulli edges through a symmetric matrix of class-pair probabilities and class proportions.The parameters are reparameterized as log odds for class membership and edge existence; asymptotic results hold up to class permutations.
5 Simulation studies
Simulations illustrate the predicted consistency, asymptotic normality, underdispersion, and computational speed of VB relative to HMC in Poisson GLMMs and LDA.
- Overall findings: VB posteriors approach the truth as sample size increases in both simulation settings, while remaining underdispersed relative to HMC.For LDA, fitted topic distributions become close to the truth by M = 1000 documents.
- Bayesian generalized linear mixed models: In the Poisson GLMM, all VB posteriors converge to their true values and exhibit approximately normal boxplot shapes.The simulation reproduces the paper’s theoretical consistency and asymptotic normality conclusions.
- Bayesian generalized linear mixed models: β1 posteriors center on their true values for N ≥1000, while the intercept β0 remains displaced until N = 20000.Continuous-variable slopes generally converge faster than binary-variable slopes, and σ2 also converges quickly.
- Computational comparison: VB takes orders of magnitude less computation time than HMC, with comparable posterior performance at N = 20000 in the GLMM simulation.In the LDA study, VB optimization converges within 10,000 steps.
- Latent Dirichlet allocation: In LDA, mean KL divergences for all K = 10 topics decrease toward zero as M increases, faster for VB than HMC.VB posterior topic distributions are very close to the truth at M = 1000 and are underdispersed compared with HMC.
6 Discussion
The discussion frames the results as theoretical support for VB while identifying broader applicability and unresolved challenges around nonparametric models, finite samples, and local optima.
- Main conclusions: The paper proves TV-consistency and asymptotic normality for VB posteriors and consistency and asymptotic normality for the VBE.The frequentist formulation assumes data arise from a fixed nonrandom true value for the global latent variables.
- Theoretical significance: The connection between ideal VB and frequentist variational approximations bridges asymptotic theory for variational estimates and VB posteriors.The authors conclude that VB remains theoretically sound despite potentially severe mean-field approximation.
- Scope: The proof framework can extend to alternative divergences and more expressive variational families containing the mean-field family.The discussion also permits misspecification when the misspecified variational log likelihood retains local asymptotic normality.
- Limitations and future work: Future work includes nonparametric VB, finite-sample properties, and theoretical characterization of local optima found by practical optimization.The paper’s asymptotic optimization analysis assumes access to the global optimum, whereas practical VB typically finds a local optimum.
A Proof of Lemma 1
The proof establishes a posterior implication by adapting an existing misspecified Bernstein–von Mises argument to dimension-dependent convergence rates.
- Proof strategy: The proof shows that the consistent testability assumption implies assumption (2.3) from Kleijn et al. (2012).This implication is needed for the paper’s posterior asymptotic analysis.
B Proof of Lemma 2
The proof controls the variational KL divergence by matching the variational family’s concentration rate to the VB ideal and handling compact-set restrictions. It then establishes consistency of the KL minimizer through these bounds and shrinking-scale arguments.
- Rate matching: The variational family must degenerate at the same rate δn as the ideal VB posterior; faster concentration makes the KL divergence diverge.The rescaled, re-centered parameter θ̌ := δn^-1(θ − µ) preserves the required rate and prevents this pathology.
- KL control: A candidate centered at θ0 has finite limiting KL divergence because both the variational distribution and ideal posterior share the same center and convergence rate.This provides the upper bound on the limiting minimum KL divergence.
- Compact restriction: Restricting the variational distribution to the compact set K does not change the limit because its scale shrinks to zero while its location remains in K.Eventually, the unrestricted and renormalized restricted distributions become arbitrarily close.
- KL control: The KL integral is bounded above and below using the LAN condition, change-of-variable calculations, and cancellation of the logdet(δn) terms.The cancellation is why the variational family must use the matching concentration rate.
- Conclusion: The proof completes the first claim by bounding the limiting KL objective and the second claim by showing compact-set restriction becomes negligible.The resulting bound establishes the stated asymptotic control of the KL minimization problem.
C Proof of Lemma 3
The proof uses Γ-convergence to compare finite-sample variational objectives with a limiting objective whose minimizers characterize the asymptotic behavior of the KL minimizer. Under moment, entropy, and tail conditions, the minimizers converge and the corresponding distributions converge in total variation.
- Γ-convergence: Γ-convergence is used because convergence of the functionals implies convergence, up to subsequences, of their minimizers to minimizers of the limiting functional.Equi-coerciveness supplies the compactness needed for convergent minimizing sequences.
- Variational family: The mean field family is parameterized by finite-dimensional parameters and uses the rescaled form q(θ) = Qd(δn^-1(θ − µ)).The change of variables centers the family at an arbitrary location µ while imposing concentration at rate δn.
- Assumptions: Assumption 2 requires continuous component densities, positive finite entropies, and conditions ensuring parameter convergence implies total variation convergence.The entropy condition prevents degenerate point-mass components and supports the continuity arguments.
- Conclusion: The minimizers converge to the limiting minimizer, and the associated variational distributions converge in total variation.The additive quadratic term is bounded and independent of the variational parameters, so it does not change the limiting argmin.
- Limiting objective: When µ ≠ θ0, the finite-sample objective diverges, whereas it remains bounded along choices centered at θ0.This separation identifies θ0 as the limiting location of the minimizers.
E Proof of Theorem 5 and Theorem 6
Theorem 5 follows by combining consistency and asymptotic normality of the KL minimizer with the fact that the VB posterior shares the VB ideal’s asymptotic properties. Theorem 6 extends the posterior-mean result to the common δn^-1 convergence rate.
- Theorem 5: Theorem 5 combines Lemmas 2–4 to establish consistency and asymptotic normality for VB posteriors.Lemmas 2 and 3 characterize the KL minimizer, while Lemma 4 transfers those properties from the VB ideal to the VB posterior.
- Theorem 6: Theorem 6 is obtained by generalizing an existing posterior-mean theorem from pn-convergence to the rate δn^-1.The proof requires replacing each occurrence of pn with δn^-1.
- Theorem 6: Assumption 3.1 supplies the condition needed for the generalized posterior-mean argument.The proof analyzes the relevant stochastic processes on compact sets and controls their approximation errors.
- Theorem 5: The rescaled variational parameter converges weakly to the KL minimizer of a normal distribution involving the limiting Gaussian experiment.This is the variational Bernstein–von Mises characterization used for the VB posterior.
F Proof of Corollary 7
The Gaussian-family proof repeats the Γ-convergence strategy for mean and covariance parameters. It shows that only mean parameters centered at θ0 retain finite limiting objective values and that the Gaussian variational distributions converge in total variation.
- Objective expansion: The proof’s objective expansion relies on Gaussian entropy, Laplace approximation, LAN, and cancellation of the two logdet(δn) terms.The covariance scaling δnΣδn is essential for this cancellation.
- Gaussian family: The Gaussian variational family uses distributions N(m, δnΣδn), which converge to a point mass as n increases.The proof evaluates the objective using Gaussian expectations and the LAN expansion.
- Γ-convergence: The Gaussian finite-sample functionals Γ-converge to a limiting functional in the mean and covariance parameters.The liminf inequality and recovery sequence are verified separately for means equal and unequal to θ0.
- Conclusion: The minimizers converge to the limiting minimizer because the remaining quadratic term is bounded and independent of m and Σ.The resulting convergence in total variation follows from a Gaussian total-variation bound.
G Proof of Lemma 8
The proof minimizes the KL divergence from a full-covariance normal distribution to a mean-field Gaussian. The optimal mean matches the target mean, while the factorized approximation matches the precision matrix at the mode.
- The KL-optimal mean-field Gaussian sets its mean equal to the full Gaussian’s mean, ˆµ0 = µ1.
- Writing Σ0 as diag(λ1,...,λd) reduces the KL minimization to optimization over its diagonal variances.
- Differentiating with respect to each λi and setting the derivatives to zero yields the optimal diagonal covariance parameters.
- Mean-field approximation matches the precision matrix at the mode.
- The entropy of the mean-field approximation is bounded above by the entropy of the full Gaussian.
H Proof of Proposition 10
The proof establishes that the variational and complete log likelihoods share the same local asymptotic normality expansion under latent-variable concentration and regularity conditions.
- The proof treats discrete local latent variables, with the continuous case adapted by constraining z near a profile point mass.
- The profile likelihood estimate zprofile is used to sandwich the variational likelihood between bounds from a point mass and Jensen’s inequality.
- Under Condition 1, the posterior of local latent variables given the true global variables concentrates around zprofile.
- The variational log likelihood Mn(θ; x) and complete log likelihood log p(x, z | θ) have the same LAN expansion around θ = θ0.
- For the variational log likelihood Ln(µ; x), a Taylor expansion, the central limit theorem, and the strong law establish LAN for s in a compact set.
- The argument additionally requires a positive definite matrix and consistent testability, with the latter supplied by consistent estimators.
- The stated corollary follows from Theorems 5 and 6 in Section 3.
J Proof of Corollary 12
The proof verifies local asymptotic normality of variational log likelihoods for generalized linear mixed and stochastic block models. It combines Taylor expansions, moment calculations, likelihood bounds, and existing asymptotic results.
- Generalized linear mixed models: The generalized linear mixed-model proof begins by verifying local asymptotic normality of the variational log likelihood.
- Generalized linear mixed models: Taylor expansions around the true parameter values and the variational frequentist estimate organize the likelihood derivative calculations.
- Generalized linear mixed models: Weak laws of large numbers, Slutsky’s theorem, and cited moment identities control the asymptotic derivative terms.
- Generalized linear mixed models: The calculation yields the full local asymptotic expansion of ℓ(β0,β1,σ2).
- Generalized linear mixed models: Consistent testability follows from the existence of consistent estimators, and the corollary follows from Theorems 5 and 6 in Section 3.
- Stochastic block models: For the stochastic block model, the variational log likelihood is defined by maximizing over q(z) in the variational family.
- Stochastic block models: A point-mass lower bound and Jensen upper bound sandwich the variational likelihood, while cited stochastic block-model lemmas characterize the asymptotics.
- Stochastic block models: The resulting LAN expansion involves asymptotically normal variables Y1 and Y2 with zero means and covariances Σ1 and Σ2.