Source-linked AI summary
Generalized Beta Mixtures of Gaussians
Artin Armagan, David B. Dunson, Merlise Clyde
TL;DR
Massive regression problems need shrinkage priors that improve on traditional choices while remaining computationally tractable, especially when variable-selection methods scale exponentially with predictors. The paper introduces generalized-beta normal scale mixtures with conjugate hierarchies and variational Bayes approximations, showing an encompassing framework for existing priors and scalable inference. Its modeling scope assumes a Gaussian linear regression setting and emphasizes prior choices that support sparse estimation and heavy-tailed behavior.
Problem
Massive-dimensional regression motivates shrinkage priors beyond normal and double exponential choices, while variable-selection approaches have computational complexity exponential in the number of predictors.
Method
The paper introduces three-parameter beta normal scale mixtures, equivalent conjugate hierarchies, and variational Bayes approximations for scalable inference.
Results
The proposed prior class encompasses many existing priors, reveals connections among their hierarchical formulations, and supports computation for massive data sets.
Takeaways & Limitations
Adjusting the global shrinkage parameter can control the number of non-zero parameters when the underlying structure is sparse.
Takeaways & Limitations
The inference formulation is presented for a Gaussian linear regression model with independent normally distributed residuals.
Abstract
from arXiv · showhide
In recent years, a rich variety of shrinkage priors have been proposed that have great promise in addressing massive regression problems. In general, these new priors can be expressed as scale mixtures of normals, but have more complex forms and better properties than traditional Cauchy and double exponential priors. We first propose a new class of normal scale mixtures through a novel generalized beta distribution that encompasses many interesting priors as special cases. This encompassing framework should prove useful in comparing competing priors, considering properties and revealing close connections. We then develop a class of variational Bayes approximations through the new hierarchy presented that will scale more efficiently to the types of truly massive data sets that are now encountered routinely.
1 Introduction
Massive-dimensional Bayesian regression motivates priors that improve on simple normal and double exponential choices while remaining computationally practical. Continuous shrinkage priors and the proposed encompassing framework address sparsity, estimation, and scalable inference.
- Penalized likelihood methods can be interpreted as posterior modes, with different priors inducing different regularization penalties.
- Variable selection priors represent coefficients with a mixture of a point mass at zero and a continuous distribution for non-zero signals.
- 2^p possible predictor subsets make Bayesian model selection and averaging computationally exponential in the number of predictors.
- Continuous shrinkage priors provide computationally attractive alternatives, including conjugate block updates that can improve MCMC mixing and convergence.
- The proposed prior class encompasses many priors, reveals hierarchical connections, and supports variational Bayes approximations designed for truly massive data sets.
2 Background
The background contrasts continuous shrinkage priors with variable-selection approaches and highlights their tail behavior and concentration near zero. Horseshoe-type priors provide strong shrinkage of noise while limiting bias for large signals.
- Heavy-tailed continuous shrinkage priors can bias large signals less while shrinking noise-like signals strongly toward zero.
- The considered prior hierarchies use normal and half-Cauchy components, with an equivalent representation obtained through ρ_j = 1/(1 + τ_j).
- The horseshoe prior uses beta-distributed shrinkage coefficients with B(1/2, 1/2), combining heavy tails with an unbounded density near zero.
- The shrinkage coefficients determine how strongly the corresponding parameters are pulled toward zero.
- NEG and NG priors offer conjugate constructions but provide less direct intuition about density behavior near the origin and in the tails.
3 Equivalence of Hierarchies via a Generalized Beta Distribution
The paper introduces the three-parameter beta distribution and its normal scale-mixture prior, then shows equivalent hierarchical forms that unify several shrinkage priors. These representations provide flexible shrinkage control and preserve conjugate structures useful for computation.
- TPB distribution: The three-parameter beta (TPB) distribution generalizes the beta distribution into a flexible scale-mixture framework.It is defined for a > 0, b > 0, and φ > 0, and is related to broader hypergeometric distribution families.
- Computational implications: The TPB hierarchy retains conjugate formulations that can simplify computation and enable variational Bayes approximations for massive data sets.The framework also provides a straightforward conjugate hierarchy for previously proposed priors.
- TPB normal scale mixture: The TPB normal scale mixture sets θj|ρj ∼ N(0, 1/ρj − 1) with ρj ∼ TPB(a, b, φ).The resulting marginal distribution is denoted TPBN(a, b, φ).
- Equivalent hierarchies: Proposition 1 gives equivalent gamma-mixture and inverted-beta representations for the TPB normal scale mixture.These equivalent hierarchies make the induced shrinkage flexible and support the later developments.
- Special cases and connections: The TPB family contains several established priors as special cases, including horseshoe, Strawderman-Berger, normal-exponential-gamma, and related formulations.In particular, a = 1 yields TPBN ≡ NEG, while (a, b, φ) = (1, 1/2, 1) yields TPBN ≡ SB ≡ NEG.
- Shrinkage behavior: The parameters separately control shrinkage behavior: smaller a increases concentration near zero, smaller b produces heavier tails, and decreasing φ favors stronger shrinkage while lightening tails.Figure 1 varies (a, b) across panels and considers φ from 1/10 through 10, with the lowest φ shown dashed.
4 Estimation and Posterior Inference in Regression Models
The paper places a generalized hierarchical shrinkage prior on regression coefficients and derives both fully Bayesian and deterministic variational inference procedures. It also characterizes how the prior’s parameters affect sparsity and behavior near zero and in the tails.
- Model and prior: The regression model assumes y = Xβ + ϵ with independent Gaussian errors of variance σ2.
- Model and prior: Each coefficient receives a normal prior with local scale τj, while λj and the global shrinkage parameter φ complete the hierarchy.The hierarchy is βj ∼ N(0, σ2τj), τj ∼ G(a, λj), and λj ∼ G(b, φ).
- Posterior inference: Variational Bayes approximates β, σ−2, τj, λj, φ, and ω through iteratively updated marginal moments.The deterministic updates provide a computational alternative to MCMC.
- Posterior inference: The conjugate hierarchy also enables straightforward Gibbs sampling for TPB normal scale-mixture priors, including the horseshoe and Strawderman-Berger priors.
- Sparse estimation: For sparse estimation, the shape parameter must satisfy 0 < a ≤ 1; when a > 1, coefficients cannot be shrunk exactly to zero.With b < 1, the marginal density retains heavy tails.
- Sparse estimation: The three illustrated parameter settings have similar tail behavior but differ drastically near the origin.Figure 2 compares prior densities for ρj and βj at a = 1/2, 1, and 3/2.
5 Experiments
The experiments evaluate the generalized-beta shrinkage framework through relative model error in two simulated cases and compare variational Bayes with Gibbs sampling in a 10,000-dimensional example. Variational Bayes identifies large signals, shrinks many others toward zero, and agrees with Gibbs-sampling estimates.
- Simulation experiments: The simulations measure median relative model error against 10-fold cross-validated lasso across two cases and multiple (a, b, φ) settings.The relative model error divides each procedure’s model error by the lasso model error; bootstrapping provides the displayed distributions.
- High-dimensional example: The high-dimensional example uses n = 100 observations and p = 10000 predictors, with 10 nonzero coefficients, each set to 3.The data-generating process has signal-to-noise ratio 3.16 and uses (a, b, φ) = (1, 1/2, 10^-4).
- High-dimensional example: The chosen φ gives prior probability P(ρj > 0.5) = 0.99, concentrating mass near total shrinkage to reflect the sparse n/p = 0.01 setting.Here φ is used a priori to limit the number of predictors in the resulting model relative to sample size.
- Inference comparison: Variational Bayes posterior means agree with Gibbs-sampling results while recovering larger signals and shrinking many zero coefficients toward zero.The Gibbs sampler used 100000 iterations, discarded 20000, and thinned the remainder; variational Bayes was run for comparison.
- Inference comparison: Treating φ as known through an informed guess substantially improves performance in the high-dimensional setting.The experiment contrasts an informed fixed choice of φ with treating φ as unknown.
6 Discussion
The discussion presents the hierarchical prior as an encompassing framework for connecting normal scale mixtures while preserving computational tractability and sparse-estimation properties. It recommends parameter choices that combine a kink at zero with heavy tails, and fixing φ when sparsity information is available or p >> n.
- Framework and computation: The hierarchical formulation connects different normal scale mixtures under a broader family and makes computation easier through equivalent hierarchies.The discussion identifies the framework as useful for understanding shrinkage behavior and recommends it over lasso for estimation performance.
- Hyperparameter choices: The recommended ranges a ∈ (0, 1] and b ∈ (0, 1), especially (a, b) = (1/2, 1/2) or (1, 1/2), produce a kink at zero and heavy tails.The kink supports sparse estimation, while heavy tails reduce unnecessary bias in large signals; b = 1/2 yields Cauchy-like tails.
- Hyperparameter choices: When sparsity is known or p >> n, φ should be fixed at a reasonable value to encode an appropriate sparsity constraint.The recommendation follows the role of φ in relating the prior sparsity constraint to the sample size.