Source-linked AI summary

Dirichlet-Laplace priors for optimal shrinkage

Anirban Bhattacharya, Debdeep Pati, Natesh S. Pillai, David B. Dunson

arXiv:1401.5398v1math.ST

TL;DR

High-dimensional Bayesian shrinkage needs both uncertainty quantification and theoretically justified posterior behavior, but point-mass mixtures are computationally difficult and continuous priors remain poorly understood. The paper proposes Dirichlet-Laplace priors with simplex-constrained scales, establishes optimal sparse posterior concentration, and develops efficient Gibbs computation, while leaving polynomial-tail cases for future work.

  • Problem

    High-dimensional Bayesian shrinkage lacks sufficient theory for posterior concentration, while point-mass mixture priors require computationally difficult stochastic model search.

  • Method

    The paper proposes Dirichlet-Laplace priors using simplex-constrained local scales and develops an efficient data-augmented Gibbs sampler with joint scale updates.

  • Results

    Theorems establish minimax-rate posterior concentration and posterior compressibility for sparse parameters under the proposed Dirichlet-Laplace prior.

  • Takeaways & Limitations

    Dirichlet-Laplace priors provide a continuous shrinkage alternative designed to combine sparse joint behavior, theoretical optimality, and efficient posterior computation.

  • Takeaways & Limitations

    The paper does not establish the lower-bound result for global-local priors with polynomial tails, such as the horseshoe.

Abstract

from arXiv · show

Penalized regression methods, such as $L_1$ regularization, are routinely used in high-dimensional applications, and there is a rich literature on optimality properties under sparsity assumptions. In the Bayesian paradigm, sparsity is routinely induced through two-component mixture priors having a probability mass at zero, but such priors encounter daunting computational problems in high dimensions. This has motivated an amazing variety of continuous shrinkage priors, which can be expressed as global-local scale mixtures of Gaussians, facilitating computation. In sharp contrast to the frequentist literature, little is known about the properties of such priors and the convergence and concentration of the corresponding posterior distribution. In this article, we propose a new class of Dirichlet--Laplace (DL) priors, which possess optimal posterior concentration and lead to efficient posterior computation exploiting results from normalized random measure theory. Finite sample performance of Dirichlet--Laplace priors relative to alternatives is assessed in simulated and real data examples.

1 Introduction

High-dimensional methods often prioritize accurate point estimation, while uncertainty quantification and Bayesian posterior concentration remain difficult under sparsity. The paper motivates shrinkage priors with optimal posterior concentration and stronger theoretical understanding.

  • High-dimensional settings can make maximum likelihood estimation break down, motivating penalization, thresholding, and especially L1 regularization.
  • Uncertainty quantification is crucial, but asymptotic confidence regions and bootstrap procedures can fail when predictors outnumber subjects.In regression, n much less than p prevents naive reliance on asymptotic normality and can make resampling inadequate.
  • Lasso corresponds to MAP estimation under a Gaussian regression model with independent Laplace priors on coefficients.This connection motivates using the full posterior distribution rather than only its mode to characterize uncertainty.
  • The desired Bayesian analogue of penalization is posterior concentration around the true parameter at an optimal rate under sparsity.The target is posterior probability tending to one for a shrinking neighborhood whose size follows the optimal convergence rate.
  • Existing shrinkage priors had little theoretical justification in high-dimensional settings, while prior concentration and effective dimensionality remain technically difficult without exact zeros.These gaps hinder understanding of posterior convergence and the behavior of continuous shrinkage priors around sparse vectors.

2 A new class of shrinkage priors

The paper develops continuous Dirichlet-kernel and Dirichlet-Laplace priors for sparse normal means, combining concentration near zero, heavy tails, joint sparsity control, and efficient posterior computation. Their theoretical target is the minimax sparse-estimation rate, while polynomial-tail cases remain unresolved.

  • 2.1 Bayesian sparsity priors in normal means problem: The normal means model observes yi = θi + ϵi with independent standard normal noise and seeks to estimate a high-dimensional mean vector.The paper notes that its ideas generalize to high-dimensional linear and generalized linear models.
  • 2.1 Bayesian sparsity priors in normal means problem: For qn-sparse means with qn = o(n), the squared minimax l2 estimation rate is 2qn log(n/qn)(1 + o(1)).Sparsity incurs only a logarithmic ambient-dimension penalty, and suitable thresholding or penalized estimators attain this rate.
  • 2.1 Bayesian sparsity priors in normal means problem: Point-mass mixture priors separate sparsity probability from signal size and can achieve minimax-optimal posterior contraction with suitable beta and tail choices.Their posterior sampling is computationally difficult because it requires stochastic search over a very large model space.
  • 2.2 Global-local shrinkage rules: Global-local priors provide continuous, heavy-tailed shrinkage with conjugate block updates, but their joint sparsity properties and prior concentration are poorly understood.They include representations corresponding to ridge, lasso, bridge, and elastic-net posterior modes.
  • 2.3 Dirichlet-kernel priors: The proposed Dirichlet-kernel construction replaces one global scale with simplex-constrained local scales, aiming to resemble the joint sparsity behavior of two-component mixtures.The simplex constraint reduces the degrees of freedom of local scales and improves control over the number of dominant coefficients.
  • 2.3 Dirichlet-kernel priors: Marginalizing local scales yields densities singular at zero while retaining exponential tails, including a wrapped Gamma distribution when a = 1/n.The prior also assigns a gamma(λ, 1/2) prior to the global scale τ, with λ = na, learned through the posterior.
  • 2.4 Posterior computation: Dirichlet-Laplace priors support efficient posterior computation through a data-augmented Gibbs sampler with a joint update for the simplex vector.Joint sampling avoids the slow mixing and convergence associated with one-at-a-time Dirichlet updates.

3 Concentration properties of Dirchlet–Laplace priors

The section develops concentration properties of Dirichlet–Laplace priors, including marginal-density behavior, posterior compressibility, and minimax-rate contraction under sparsity conditions.

  • The joint Dirichlet–Laplace density admits analytically convenient product-marginal representations and bounds near zero and infinity.These representations support the subsequent concentration analysis and contraction proof.
  • When an = n−(1+β) for small β > 0, the posterior contracts at the minimax rate under qn = o(n) and ∥θ0∥2^2 ≤ qn log^4 n.The rate is expressed through s_n^2 = qn log(n/qn), and the conclusion also holds for an = 1/n under an additional sparsity condition.
  • The paper presents this as the first posterior-contraction result for a continuous shrinkage prior in the normal-means or closely related high-dimensional regression setting.It contrasts the result with a reported suboptimal rate for a broad subclass of global-local priors, including the Bayesian lasso.
  • The stated lower-bound result does not cover global-local priors with polynomial tails, such as the horseshoe.The authors leave that case for future work and conjecture broader optimality based on empirical performance.
  • The Dirichlet–Laplace prior has a marginal density with an infinite spike near zero, concentrating substantial prior mass around sparse signals.The section quantifies small-neighborhood probability and uses this near-zero behavior in posterior-contraction arguments.
  • For continuous shrinkage priors, approximate model size is measured by |suppδ(θ)|, the number of coefficients exceeding δ in magnitude.This measure replaces exact support size because continuous priors assign probability zero to θ = 0.
  • 3.1 Auxiliary results: Posterior compressibility ensures that approximate dimensionality does not substantially exceed the true sparsity level qn.For δn = qn/n and an = n−(1+β), the relevant support-size bound holds; with an = 1/n, it holds when qn ≳ log n.

4 Simulation Study

The simulation study evaluates Dirichlet–Laplace priors across dimensions and sparsity levels, comparing posterior-median squared error with several shrinkage alternatives. DL1/n performs strongly overall, while larger a can reduce excessive shrinkage of smaller signals and offer computational gains.

  • Simulation design: The study uses 100 replicates of sparse normal-means data across n = 100, 200, sparsity levels of 5%, 10%, and 20%, and signal sizes A = 7, 8.Each vector has qn non-zero entries, all equal to A.
  • Simulation design: Table 1 reports average posterior-median squared error for Bayesian lasso, Dirichlet–Laplace, Lasso, EBMed, point-mass, and horseshoe methods.The comparison includes both continuous shrinkage and point-mass approaches.
  • Main results: DL1/n shows a wide performance advantage over Bayesian lasso in Table 1, while horseshoe performs similarly to DL1/n.The authors attribute DL1/n’s superior performance to strong concentration around zero.
  • Main results: DL1/n can shrink several relatively small signals toward zero, so a larger a can soften the singularity at zero when that trade-off is preferable.For a = 1/2, computational gains arise because the distribution of Tj has an inverse-Gaussian form with exact samplers available.
  • Prior sensitivity: A second simulation varies A from 2 to 7 in a 1000-dimensional problem containing large, moderate, and null signals to examine prior elicitation needs.Table 2 summarizes squared error averaged over 100 replicates for Bayesian shrinkage priors.
  • Visualization: Figures 2 and 3 visualize observations, posterior medians, and 95% pointwise credible intervals for Bayesian lasso, DL1/n, horseshoe, and DL1/2.The displayed setting is n = 200, qn = 10, and A = 7.

5 Prostate data application

The prostate-cancer application analyzes gene-expression differences between 50 controls and 52 patients across 6033 genes. A fully Bayesian DL analysis selects 128 non-null genes, overlapping substantially with empirical-Bayes and FDR selections while being less conservative than horseshoe.

  • Data and goal: The dataset contains expression levels for 6033 genes measured on 102 subjects: 50 controls and 52 prostate-cancer patients.The goal is to identify genes whose expression differs between the two groups.
  • Method: The analysis uses a normal-means model zi = θi + ϵi and assigns θ a DLa prior.The prior parameter a is selected rather than fixed.
  • Computation: The Gibbs sampler retained 5000 samples after 5000 burn-in draws, with average effective sample size 2369.2 across θi and approximately linear scaling per iteration.The posterior mode of a was 1/20.
  • Selection: The selection procedure clusters |θi| into two groups at each MCMC iteration and estimates the number of signals from the smaller cluster.This operationalizes the expected separation between near-zero and nonzero effects.
  • Results: The DL method declares 128 genes non-null; 100 overlap with EBMed, and all 54 FDR-selected genes are included.Horseshoe selects only one gene under the same clustering procedure.

6 Proof of Theorem 3.1

The proof establishes posterior concentration by controlling prior-mass ratios and testing deviations over sparse support sets. A covering-net argument bounds the relevant regions and yields convergence outside a shrinking Euclidean ball.

  • Setup: The proof treats an = 1/n in detail and indicates how the argument changes for an = n−(1+β).The latter choice removes the requirement qn ≳ log n.
  • Setup: The argument defines δn = rn/n and restricts attention to support sets S with |S| ≤ Aqn.Approximate support is determined by coordinates exceeding δn in magnitude.
  • Covering argument: For each sparse support and distance shell, a 2jrn-net is constructed, and balls centered at its net points cover the corresponding parameter region.The net size is controlled exponentially in the support size.
  • Testing argument: The testing step reduces posterior control to bounding βS,j,i, defined as a probability ratio involving a covering ball and a neighborhood of θ0.The proof combines this ratio with prior concentration and shell-wise testing bounds.
  • Key bound: Lemma 6.1 bounds log βS,j,i by |S| log(2j) + C(|S| + |S0|) log n + C′r2n.The same conclusion is used in the subsequent posterior bound.
  • Conclusion: The final convergence statement follows from the preceding bounds and the theorem’s restriction to sufficiently sparse approximate supports.The displayed result is obtained after combining the exponential bounds with the testing construction.

Proof of Proposition 2.1

The proposition derives the marginal density of a coefficient by integrating over its Beta-distributed shrinkage weight. In the special and general parameterizations, the resulting density diverges at zero, establishing an infinite spike.

  • Special case: When a = 1/n, the marginal shrinkage weight satisfies φj ∼ Beta(1/n, 1 − 1/n), and the coefficient density is obtained by integrating over φj.The resulting integral is transformed using z = φj/(1 − φj).
  • General case: For general a, the marginal weight satisfies φj ∼ Beta(a, (n − 1)a), leading to an analogous integral representation for θj.The same change of variables is used in the derivation.
  • Singularity at zero: The transformed integral admits a lower bound that diverges as |θj| → 0, proving that the marginal density has an infinite spike at the origin.This behavior follows by monotone convergence.

Proof of Theorem 2.2

The proof derives the joint posterior through normalized random-measure theory by normalizing independent variables and matching the resulting density to the target expression. Comparing exponents identifies δ = 2 − a and links the component density to a generalized inverse Gaussian distribution.

  • Normalized random measures: The proof invokes a normalized-random-measure result for normalized independent positive random variables.The variables are normalized by their sum, yielding a joint density on the simplex.
  • Density construction: Setting fj(x) proportional to x^-δ exp(−|θj|/x) exp(−x/2) produces the density used in the posterior derivation.
  • Parameter matching: Comparing exponents gives δ = 2 − a, while the remaining requirement holds because λ = na.
  • Parameter matching: When δ = 2 − a, fj corresponds to a giG(a − 1, 1, 2|θj|) distribution.

Proof of Proposition 3.1

The proof concludes by applying the cited result 8.432.7 from reference.

  • The proposition follows directly from result 8.432.7 in reference.
  • The cited result is used as the final justification for the proposition.
  • No additional derivation is given in this proof passage beyond invoking the cited result.

Proof of Lemma 3.2

The lemma bounds prior probabilities using monotonicity of h and asymptotic bounds for modified Bessel functions, followed by repeated Cauchy–Schwarz inequalities.

  • Monotonicity of Π(x), and hence h(x), in |x| bounds log ΠS(η) by |S|h(δ).
  • For small arguments, Bessel and gamma-function asymptotics imply Π(δ) is proportional to a^-1|δ|^(a−1).
  • The resulting small-δ bound is h(δ) proportional to (1−a) log(δ^-1) − log a^-1 plus a constant, and is bounded by C log(δ^-1).
  • For large |x|, a lower bound for Kα(z) yields an upper bound on −h(x) involving log a^-1 and 3/2 log |x|.
  • The proof then applies the Cauchy–Schwarz inequality twice.

Proof of Lemma 3.3

The proof controls posterior concentration by bounding prior mass outside sparse supports and small inactive coefficients, with the resulting error probabilities tending to zero under logarithmic sparsity conditions.

  • Using the representation of θ1 conditional on ψ1, the tail probability satisfies P(|θ1| > δ | ψ1) = e^-δ/ψ1.
  • An incomplete-gamma bound supplies a small-δ probability estimate, completed using boundedness of (1/2)a and a logarithmic inequality.
  • The proof reduces the required Gaussian probability bound by defining S0 = supp(θ0) and using |S0| = qn.
  • Concentration conclusion: The posterior probability of the remaining support event converges to zero under the stated conditioning.
  • Inactive coordinates: The inactive-coordinate probability is bounded using independence and P(|θ1| < rn/√n) ≥ 1 − log n/n.
  • Support control: The support-size probability is controlled with a binomial model for |suppδn(θ)| and a Chernoff inequality.
  • Support control: For an = 1/n and qn ≥ C0 log n, choosing an = Aqn/n gives ζn < an for suitable A.
  • Concentration conclusion: Choosing A sufficiently large makes the relevant probability bounds vanish, including an e^-Aqn log 2 bound and a final expression tending to zero.
Loading 1401.5398v1…