Source-linked AI summary

Bayesian variable selection with shrinking and diffusing priors

Naveen Naidu Narisetty, Xuming He

arXiv:1405.6545v1math.ST

TL;DR

High-dimensional variable selection requires a Bayesian procedure that can identify sparse models when the number of covariates is large relative to the sample size. The paper uses sample-size-dependent shrinking and diffusing Gaussian spike-and-slab priors with a model-space prior, and shows that the posterior probability of the true model converges to one while enabling Gibbs sampling. It also establishes an asymptotic relationship with L0-penalized model selection and reports good performance in simulations and a real-data example.

  • Problem

    The paper addresses Bayesian variable selection for sparse high-dimensional regression, where p can be large relative to n and selection consistency is needed.

  • Method

    The method uses sample-size-dependent shrinking and diffusing Gaussian spike-and-slab priors together with a prior on the model space, sampled using a standard Gibbs sampler.

  • Results

    The posterior probability of the true model converges to one, including when p = e^o(n), and the method is asymptotically related to L0-penalized likelihood selection.

  • Takeaways & Limitations

    The approach combines strong selection consistency with transparent prior tuning and computationally convenient posterior sampling.

  • Takeaways & Limitations

    The theoretical conditions restrict the number of covariates to be no greater than exponential in n and include regularity conditions that help identify the true model.

Abstract

from arXiv · show

We consider a Bayesian approach to variable selection in the presence of high dimensional covariates based on a hierarchical model that places prior distributions on the regression coefficients as well as on the model space. We adopt the well-known spike and slab Gaussian priors with a distinct feature, that is, the prior variances depend on the sample size through which appropriate shrinkage can be achieved. We show the strong selection consistency of the proposed method in the sense that the posterior probability of the true model converges to one even when the number of covariates grows nearly exponentially with the sample size. This is arguably the strongest selection consistency result that has been available in the Bayesian variable selection literature; yet the proposed method can be carried out through posterior sampling with a simple Gibbs sampler. Furthermore, we argue that the proposed method is asymptotically similar to model selection with the $L_0$ penalty. We also demonstrate through empirical work the fine performance of the proposed approach relative to some state of the art alternatives.

1. Introduction.

The paper develops a Bayesian method for selecting sparse active covariates in high-dimensional linear regression. Its shrinking and diffusing spike-and-slab priors support strong selection consistency while remaining computationally convenient.

  • High-dimensional regression can be ill-posed when p > n, motivating selection of the small set of covariates with nonzero coefficients.
  • Bayesian variable-selection methods combine latent activity indicators with concentrated spike and diffuse slab priors for regression coefficients.The binary variable Zi indicates whether covariate i is active.
  • When p diverges, setting marginal inclusion probabilities of order p^-1 makes models with diverging size receive vanishing prior probability.This acts as prior penalization against unnecessarily large models.
  • The proposed shrinking and diffusing priors achieve strong selection consistency for p = e^o(n), and posterior sampling can use a standard Gibbs sampler.
  • The resulting model-space selection is asymptotically closely related to L0-penalized likelihood selection.

3. Orthogonal design.

The orthogonal-design analysis explains why prior parameters must depend on sample size. Shrinking and diffusing specifications can recover active and inactive covariates marginally, whereas fixed choices can fail as dimensionality grows.

  • 3. Orthogonal design: Under an orthogonal design, the joint posterior factorizes across covariates, making each latent activity indicator conditionally independent of the others given the data.The marginal posterior of Zi can therefore be analyzed separately.
  • 3.1. Fixed parameters: Fixed prior parameters can fail to identify an active coefficient: its limiting inclusion probability remains below 0.5 with high probability.
  • 3.2. Shrinking τ^2_0,n: With a shrinking spike variance and fixed slab variance and inclusion probability, inactive covariates are excluded and active covariates are included with posterior probability converging to one.
  • 3. Orthogonal design: Marginal convergence of each Zi does not by itself guarantee consistency of the overall selected model.Overall model consistency requires sample-size-dependent prior parameters.
  • 3.3. Shrinking and diffusing priors: With fixed slab and inclusion parameters, selection becomes inconsistent when the number of covariates is much greater than √n; diffusing priors replace the threshold σ2 log n/n with (2 + δ)σ2 log pn/n.

4. Main results.

The paper establishes strong selection consistency for its Bayesian variable-selection method under high-dimensional conditions, including nearly exponential covariate growth. The results also support computation using marginal posterior probabilities and accommodate a prior on the variance under a model-size restriction.

  • Conditions: The theory assumes p_n = e^{d_n n} with d_n → 0, equivalently log p_n = o(n), while the true model size remains fixed.The conditions also impose identifiability and regularity requirements on the design and signal strength.
  • Posterior convergence: P(Z = t | Y, σ^2) → 1 as n → ∞, so the posterior probability of the true model converges to one.This is the paper’s strong selection-consistency result for known σ^2 under Conditions 4.1–4.5.
  • Variance specification: Strong selection consistency also holds for misspecified variance ˜σ^2 when the stated inequalities relating ∆_n, γ_n, σ^2, and ˜σ^2 are satisfied.Thus, the theorem does not require the true variance to be known exactly.
  • Posterior convergence: Posterior ratios are controlled across unrealistically large, over-fitted, large, and under-fitted model classes.The proof separately analyzes models that exceed the dimension threshold, add inactive covariates, omit active covariates, or have moderate dimension while missing active covariates.
  • Variance specification: With an inverse Gamma prior on σ^2, P(Z = t | Y) → 1 when models are restricted to dimension at most |t| + wn/log p_n.The excluded models have dimension of order n/log p_n and are characterized as unrealistically large.
  • Marginal selection: Marginal posterior probabilities can identify the active covariates with probability tending to one, providing a computationally useful variable-selection rule.The corollary is presented as a direct consequence of the two main posterior-convergence theorems.

5. Connection with penalization methods.

The paper connects its Bayesian model-selection procedure to an L0-penalized objective through the asymptotic behavior of its MAP estimate. Its penalty treats inactive covariates differently from magnitude-based L1 and SCAD penalties.

  • MAP and penalization: The MAP estimate is asymptotically described by minimizing an objective function derived from the Bayesian setup.This follows from the posterior-ratio bounds and the asymptotic behavior of the prior parameters.
  • MAP and penalization: The resulting model-space selection is closely related to L0-penalized likelihood.The paper states that its selection properties are similar to those of L0 penalization.
  • Penalty behavior: Inactive covariates receive a penalty of order log(n ∨ p_n), irrespective of coefficient magnitude.This contrasts with L1 and SCAD penalties, which are proportional to coefficient magnitude over an interval near zero.
  • Penalty behavior: AIC and BIC are identified as special cases of L0 penalization, with penalty quotients 2 and log n, respectively.The paper compares these criteria with its objective through the corresponding penalty term ψ_n,k.

6. Discussion of the conditions.

The conditions permit nearly exponentially many covariates and relatively broad design matrices while controlling correlations, signal strength, and Gram-matrix eigenvalues. Sub-Gaussian random designs satisfy a stronger minimum-eigenvalue property with high probability.

  • Growth and prior rates: The framework allows the number of covariates to grow exponentially with n and permits regression coefficients to depend on n.Condition 4.1 restricts covariates to no more than exponential in n; Conditions 4.3–4.5 allow β to depend on n.
  • Growth and prior rates: The prior conditions specify shrinking and diffusing rates for the spike and slab variances.Condition 4.2 links the prior parameters to these rates, while Condition 4.5 relates Gram-matrix eigenvalues to the prior parameters.
  • Identification conditions: Condition 4.4 restricts active–inactive correlation and imposes a lower bound on the signal-to-noise ratio needed to identify the true model.The condition is described as a mild regularity requirement for model identification.
  • Design flexibility: Independent isotropic sub-Gaussian rows yield uniformly positive minimum eigenvalues for all sufficiently small submatrices with high probability.Lemma 6.1 establishes this property, which is stronger than Condition 4.5.
  • Design flexibility: Unlike restricted isometry conditions, Condition 4.5 allows the minimum eigenvalue to be zero, including perfect correlation among inactive or active covariates.This makes the admissible design class wider than conditions requiring all relevant eigenvalues to be bounded away from zero.

7. Computation.

The proposed sampler uses conjugate full conditionals to draw posterior samples of the latent inclusion indicators and coefficients. High-dimensional coefficient updates remain tractable through block updating.

  • Gibbs sampling: A standard Gibbs sampler draws posterior samples of the latent inclusion vector Z using conjugate full conditionals.The coefficient, inclusion-indicator, and variance conditionals have standard forms.
  • Gibbs sampling: The conditional coefficient distribution is multivariate normal with covariance structure (X′X + D_k)^−1 and mean V X′Y.Here V = (X′X + D_k)^−1.
  • Gibbs sampling: The conditional variance distribution is inverse Gamma, with parameters incorporating the prior quadratic form and regression residual sum of squares.Its shape includes α1, n/2, and p_n/2; its scale includes β′D_kβ/2 and residual error.
  • Computational scalability: Block updating efficiently samples the high-dimensional coefficient vector by drawing from smaller-dimensional normal distributions.This addresses the main computational difficulty when p_n is large.

8. Simulation study.

The simulations compare BASAD with Bayesian and penalization methods across dimensionality, signal strength, correlation, and sparsity settings. BASAD generally performs strongly for model selection and false-discovery control, while prediction remains competitive and tuning matters for difficult regimes.

  • Experimental design: BASAD is evaluated across six cases varying n, p, correlations, signal strengths, and sparsity levels.The study includes both p ≥ n and n > p settings, with repeated simulated datasets.
  • Experimental design: The reported metrics include posterior probabilities for inactive and active covariates, exact-model selection, active-set inclusion, FDR, and MSPE.BASAD denotes the median probability model, while BASAD.BIC uses a BIC-selected threshold.
  • Main findings: BASAD and piMOM generally select the true model more effectively and control false discoveries better than the other methods, with BASAD standing out.Penalization methods more often include all active covariates but incur overfitting and false discoveries.
  • Main findings: BASAD remains competitive in prediction error but does not always outperform its competitors.The paper distinguishes model-selection performance from prediction accuracy.
  • Main findings: When signals are low, all methods struggle to recover the correct model; BASAD.BIC has lower prediction error but slightly higher false-positive rates in most cases.The two BASAD variants have similar prediction errors in this setting.
  • Correlation and sparsity: Under moderate inactive and active–inactive correlations, BASAD outperforms the alternatives in case 4.The authors relate this performance to BASAD’s similarity to the L0 penalty and its ability to accommodate such correlations.
  • Correlation and sparsity: With 25 active covariates, the default K = 10 performs poorly, whereas K = 50 considerably improves BASAD’s performance.The authors recommend a generous K or BIC to choose among K values when sparsity is uncertain.

9. Real data example.

The real-data analysis evaluates variable-selection methods on gene-expression data for PEPCK and GPAT after marginal screening. BASAD generally achieves lower prediction error across model sizes and selects less correlated variables.

  • The dataset contains 22,575 genes measured on 60 mice, with PEPCK and GPAT as response phenotypes.
  • Marginal-correlation screening reduces each response dataset to p = 200 or p = 400 predictors before model selection.The screened predictors include the intercept, gender, and 198 or 398 genes.
  • The analysis compares BASAD with LASSO, SCAD, and BCR using repeated training-test splits and average MSPE.Each repetition trains on 55 observations and tests on five; the process is repeated 100 times.
  • BASAD's MSPE is mostly smaller than those of the other methods across model sizes for both responses.The figure reports MSPE versus model size for PEPCK and GPAT at p = 200 and p = 400.

10. Conclusion.

The conclusion presents shrinking and diffusing spike-and-slab priors as a consistent Bayesian variable-selection method for high-dimensional data. It also emphasizes computational simplicity, an L0-penalty connection, and strong empirical performance, while noting that prediction is not the paper's primary focus.

  • The proposed method achieves strong selection consistency, with posterior probability of the true model converging to one under mild conditions.
  • The method uses transparent prior-tuning parameters and supports posterior sampling with a standard Gibbs sampler.
  • The proposed approach has an asymptotic relationship with L0-penalty model selection and performs well in simulations and real-data analysis.The empirical studies use default tuning parameters without attempting to optimize them.
  • Selection consistency is established under Gaussian-error analysis, although the authors state that Gaussian errors are not necessary in principle.Other error distributions would require suitable deviation inequalities for quadratic forms.
  • The paper's primary focus is model-selection consistency rather than prediction accuracy.Prediction can improve in low-signal cases through BIC or cross-validation tuning, or through model averaging.

11. Proofs.

The proofs establish posterior properties through matrix identities, shrinkage representations, and concentration arguments for high-dimensional design matrices. The supplied proof passages include both posterior algebra and sub-Gaussian random-design bounds.

  • The appendix proves selected lemmas and refers to Narisetty and He (2014) for proofs of remaining results.
  • The joint posterior of β, σ^2, and Z is used as the starting point for proving Lemma 4.1.
  • The posterior expression includes inverse-gamma hyperparameters, model size, and prior-variance terms before algebraic rearrangement.
  • The estimator β̃ = (D_k + X′X)^−1X′Y shrinks coefficients excluded from model k toward zero while leaving included coefficients nearly unshrunk.The shrinkage is governed by D_k, the precision matrix of β conditional on the model indicators.
  • The proof uses the Sherman–Morrison–Woodbury identity, determinant identities, and eigenvalue bounds for Gram matrices.
  • For sub-Gaussian isotropic design rows, concentration inequalities and a union bound control the relevant events across candidate models.The argument assumes model sizes satisfy |k| ≤ m_n with |k| = o(n).

Supplement to “Bayesian variable selection with shrinking and diffusing

The supplement contains proofs for several main theoretical results of the paper. It serves as a technical companion to the published analysis.

  • The supplement is associated with the paper on Bayesian variable selection with shrinking and diffusing priors.
  • It contains proofs of Theorems 4.1 and 4.2.
  • It also contains the proof of Lemma 4.2.
Loading 1405.6545v1…