Source-linked AI summary
Bayesian Variable Selection and Estimation for Group Lasso
Xiaofan Xu, Malay Ghosh
TL;DR
The paper addresses how to perform variable selection and estimation when predictors are naturally grouped, including selection within groups. It develops spike-and-slab Bayesian group and sparse-group models linked to penalized regression, and finds that posterior median thresholding provides strong selection and estimation performance, subject to an estimation-efficiency limitation under a stated penalty condition.
Problem
Existing group-selection methods do not uniformly address selection at both group and individual levels, while posterior-mean Bayesian estimators do not produce exact zero estimates.
Method
The paper develops Bayesian group and sparse-group selection models using spike-and-slab priors, posterior median thresholding, and Gibbs-sampling-based inference.
Results
Posterior median thresholding achieves strong variable-selection accuracy and prediction performance, with at least as good and sometimes better true- and false-positive rates than the highest posterior probability model.
Takeaways & Limitations
Posterior median estimators offer a sparse Bayesian alternative that can select and estimate simultaneously at group and within-group levels.
Takeaways & Limitations
Under the examined orthogonal-design setting, the penalty condition that supports selection consistency does not provide the optimal estimation rate.
Abstract
from arXiv · showhide
The paper revisits the Bayesian group lasso and uses spike and slab priors for group variable selection. In the process, the connection of our model with penalized regression is demonstrated, and the role of posterior median for thresholding is pointed out. We show that the posterior median estimator has the oracle property for group variable selection and estimation under orthogonal designs, while the group lasso has suboptimal asymptotic estimation rate when variable selection consistency is achieved. Next we consider bi-level selection problem and propose the Bayesian sparse group selection again with spike and slab priors to select variables both at the group level and also within a group. We demonstrate via simulation that the posterior median estimator of our spike and slab models has excellent performance for both variable selection and estimation.
1 Introduction
The paper develops Bayesian spike-and-slab methods for selecting predictors at group and within-group levels. It connects these models to penalized regression and studies posterior median thresholding as a joint selection-and-estimation procedure.
- Motivation: Grouped predictors arise in categorical, nonlinear, genomic, and domain-informed regression settings, motivating group-wise variable selection.Examples include dummy variables, basis functions, and genes in the same biological pathway.
- Group selection: The Bayesian group lasso with spike-and-slab priors produces exact zero estimates at the group level for group variable selection.The model uses a multivariate point-mass mixture prior for each predictor group.
- Empirical comparison: The proposed Bayesian models are designed to retain prediction accuracy while improving variable-selection performance relative to group and sparse group lasso methods.The introduction reports comparable prediction accuracy for group selection and significant simulation improvements for bi-level selection.
- Bi-level selection: For bi-level selection, the Bayesian sparse group selection model uses hierarchical spike-and-slab priors to select groups and individual variables within groups.It is proposed as an alternative to sparse group lasso and related methods that use posterior credible intervals for within-group selection.
- Posterior median thresholding: Posterior median thresholding can select and estimate simultaneously, with oracle properties under orthogonal designs and lower false-positive rates than lasso-based methods in simulations.The paper contrasts this with the group lasso, which sacrifices estimation rate to achieve selection consistency.
- Bayesian computation: The paper derives Gibbs samplers and posterior mean and median estimators for its fully Bayesian group-selection formulation.Posterior median thresholding is introduced and its frequentist oracle property is proved for orthogonal designs.
2 Bayesian Group Lasso with Spike and Slab Prior (BGL-SS)
BGL-SS combines a group-level point-mass spike with a Multi-Laplace slab, yielding exact zero group estimates while shrinking coefficients in selected groups. Its posterior median acts as a thresholding estimator, with oracle properties under orthogonal designs and a rate advantage over group lasso when selection consistency is required.
- Model formulation: BGL-SS uses a multivariate zero-inflated mixture prior to introduce exact sparsity at the group level.The prior places point mass at zero alongside a group-lasso-type slab.
- Marginal prior: The prior combines a point mass that produces exact zero coefficients with group-level shrinkage for selected groups.These two components jointly support group selection and estimation.
- Connection with penalized regression: Under the reparameterization βg = γgbg, the posterior mode corresponds to penalized regression with an L2 group penalty and an L0-like penalty on nonzero groups.The L0-like component makes direct optimization combinatorial for moderate or large numbers of groups.
- Posterior median thresholding: With block-orthogonal designs, the marginal posterior median sets an entire group to zero when its block least-squares norm falls below a threshold.The threshold depends on the prior configuration, including π0.
- Asymptotic properties: Under orthogonal designs, median thresholding has variable-selection consistency and asymptotic normality, whereas group lasso selection consistency requires a tuning regime that yields a slower estimation rate.The paper identifies the group lasso convergence rate as n/λn and contrasts it with the median estimator’s oracle property.
3 Bi-level Selection
The paper develops Bayesian methods for selecting variables at both the group and within-group levels, building from sparse group lasso representations to spike-and-slab models that produce sparsity at both levels.
- The sparse group lasso combines group-wise and within-group sparsity, and its estimator is equivalent to a MAP solution under the corresponding prior.
- The Bayesian sparse group lasso (BSGL) uses a two-level hierarchical model to shrink coefficients both across groups and within groups.Its hierarchical construction uses Gaussian priors with parameters controlling the two shrinkage levels.
- Because BSGL posterior mean and median estimators are never exactly zero, the paper proposes BSGS-SS to achieve sparsity at both selection levels.The proposed model uses spike-and-slab priors for group and individual variable selection.
- BSGS-SS separates group-level and within-group sparsity through a coefficient reparameterization and corresponding spike-and-slab priors.Group coefficients receive multivariate spike-and-slab priors, while nonnegative scale parameters receive individual-level spike-and-slab priors.
- An alternative binary-masking formulation separately represents group and individual inclusion indicators and is expected to have comparable performance to BSGS-SS.
4 Simulation
Simulations compare spike-and-slab Bayesian estimators with group-lasso methods across five examples, evaluating selection accuracy, prediction error, prior sensitivity, and posterior mean versus median estimation.
- Simulation designs: The simulations compare BGL-SS, BSGS-SS, and related lasso methods across five examples with group-level and within-group sparsity settings.Examples include low-dimensional, high-dimensional, correlated-predictor, polynomial, and categorical-factor designs.
- Variable selection: Median thresholding models are more parsimonious and outperform the compared methods, including highest posterior probability models, in model-selection accuracy.The comparison uses true-positive and false-positive rates over 50 simulations.
- Prediction: BGL-SS has prediction error comparable to group lasso except in Example 2, while BSGS-SS outperforms sparse group lasso in all examples.Prediction errors are summarized by median mean squared error across 50 replications.
- Estimator comparison: Posterior mean and posterior median estimators have very close prediction errors, but median thresholding provides stronger sparse selection.In Example 1 variants, posterior medians identify the two most important factors, whereas posterior means can lack sufficient shrinkage under high noise.
- Estimator comparison: BGL-SS is favored when relevant groups lack within-group sparsity, whereas BSGS-SS improves prediction when substantial within-group sparsity is present.BSGS-SS can identify within-group sparsity, though it may shrink a very small true coefficient.
- Sensitivity: A flat prior with mean 1/2 on π0 performs poorly for BGL-SS in high-dimensional settings where most predictor groups are zero.The same prior still yields better variable selection than group lasso in the reported setting.
5 Discussion
The discussion explains how spike-and-slab priors produce sparse group estimators and why posterior median thresholding is useful for simultaneous selection and estimation.
- Sparse estimation: Traditional Bayesian group-lasso scale-mixture priors shrink group coefficients but do not produce sparse estimators.Spike-and-slab priors instead place point mass at zero, including multivariate point mass for coefficient groups.
- Posterior median: The posterior median estimator automatically performs both selection and estimation, like the lasso estimator.The median thresholding rule is more parsimonious than the median probability model.
- Empirical implication: Posterior median thresholding selects fewer variables than group-lasso methods while achieving similar or sometimes better prediction error.Its true- and false-positive rates are at least as good as, and sometimes better than, those of the highest probability model.
Appendix A: Propriety of (21)
Appendix A states that prior (21) is proper.
- Prior (21) is proper.
Appendix B: Marginal Prior for The Bayesian Sparse Group Lasso
Appendix B introduces the marginal prior on βg derived from priors (20) and (21).
- The marginal prior on βg is obtained from priors (20) and (21).
Appendix C: Gibbs Sampler for BSGL
The appendix specifies the posterior density and the full conditional distributions used to generate posterior samples for the BSGL model.
- The Gibbs sampler begins by specifying the joint posterior density of β, τ, γ, and σ2 conditional on Y and X.
- Posterior sampling proceeds through the model’s full conditional posterior distributions.
- The conditional-update notation indexes groups g = 1, . . . , G and within-group variables j = 1, . . . , mg.
- The sampler includes a block-structured matrix with group-specific diagonal blocks V1, V2, and subsequent entries.