Source-linked AI summary
The Horseshoe+ Estimator of Ultra-Sparse Signals
Anindya Bhadra, Jyotishka Datta, Nicholas G. Polson, Brandon Willard
TL;DR
Ultra-sparse signal detection needs priors that extract weak signals while preserving desirable decision-theoretic properties. This paper introduces the horseshoe+ prior and shows theoretically and empirically that it improves estimation and testing performance over existing sparse estimators.
Problem
Sparse signal detection requires priors that handle negligible effects while preserving computationally feasible estimation and testing in high dimensions.
Method
The paper extends the horseshoe global-local prior with a horseshoe+ hierarchy and analyzes its sparse marginal densities using Meijer-G functions.
Results
The horseshoe+ estimator attains lower MSE, reaches Bayes-oracle testing risk up to O(1), and has a sharper type-I error bound than horseshoe.
Takeaways & Limitations
The horseshoe+ prior provides substantial improvement for nearly-black or ultra-sparse signals by better separating signals from noise and handling sparsity.
Takeaways & Limitations
The testing results concern the finite, non-zero asymptotic regime known as the verge of detectability.
Abstract
from arXiv · showhide
We propose a new prior for ultra-sparse signal detection that we term the "horseshoe+ prior." The horseshoe+ prior is a natural extension of the horseshoe prior that has achieved success in the estimation and detection of sparse signals and has been shown to possess a number of desirable theoretical properties while enjoying computational feasibility in high dimensions. The horseshoe+ prior builds upon these advantages. Our work proves that the horseshoe+ posterior concentrates at a rate faster than that of the horseshoe in the Kullback-Leibler (K-L) sense. We also establish theoretically that the proposed estimator has lower posterior mean squared error in estimating signals compared to the horseshoe and achieves the optimal Bayes risk in testing up to a constant. For global-local scale mixture priors, we develop a new technique for analyzing the marginal sparse prior densities using the class of Meijer-G functions. In simulations, the horseshoe+ estimator demonstrates superior performance in a standard design setting against competing methods, including the horseshoe and Dirichlet-Laplace estimators. We conclude with an illustration on a prostate cancer data set and by pointing out some directions for future research.
1 Introduction
The paper introduces the horseshoe+ prior as a global-local shrinkage extension designed to improve ultra-sparse signal estimation and testing. It establishes sharper theoretical guarantees than the horseshoe and reports superior simulation performance against competing shrinkage estimators.
- Motivation: The horseshoe+ prior extends the horseshoe framework to sharpen signal extraction in ultra-sparse normal-means problems while preserving desirable decision-rule properties.The motivating setting contains a sublinear number of nonzero parameters, with p_n=o(n).
- Theoretical properties: The horseshoe+ decision rule attains Bayes-oracle risk under 0-1 loss up to a multiplicative constant and has a sharper type-I error bound than the horseshoe.The constant in Bayes risk is close to the oracle constant.
- Theoretical properties: The horseshoe+ posterior mean squared error is always smaller than the horseshoe posterior mean squared error when estimating a large signal.This comparison is specifically stated for large signals.
- Theoretical properties: The horseshoe+ estimated sampling density converges super-efficiently at zero in K-L distance, with an asymptotic risk upper bound smaller than the horseshoe bound.The analysis uses Meijer-G functions to study the prior’s asymptotic properties.
- Empirical evaluation: In standard-design simulations, horseshoe+ outperforms both the horseshoe and Dirichlet-Laplace estimators under squared-error estimation and 0-1-loss testing.The paper also illustrates the proposed prior on a high-dimensional prostate cancer data set.
2 The one and two groups models
This section contrasts two-groups spike-and-slab modeling with one-group global-local shrinkage priors for ultra-sparse signal estimation and testing. It presents the two-groups model as a theoretical gold standard while motivating one-group alternatives for negligible but nonzero effects and computational simplicity.
- Two-groups model: The two-groups model produces sparse estimates with exact zeros and has established minimax estimation and asymptotically optimal testing properties.Its thresholding-based empirical Bayes and full Bayes estimators are minimax in ℓ2, while testing performance matches the Bayes oracle up to a constant.
- One-group model: One-group global-local priors are motivated when most effects are negligible but not exactly zero, as in high-dimensional, low-sample-size gene-expression studies.This provides an argument against the exact sparsity induced by the two-groups model.
- Two-groups model: In the two-groups posterior mean, global shrinkage pulls all parameters toward zero while local inclusion probabilities let signals escape excessive shrinkage.The absence of local shrinkage explains why Stein-type global shrinkage estimators perform poorly in nearly-black settings.
- One-group model: The horseshoe prior is a one-group global-local prior designed to reproduce the desirable posterior-mean behavior of the two-groups model with easier computation.It has been shown to possess attractive theoretical properties, including super-efficient K-L contraction, asymptotic Bayes-optimal multiple testing up to a constant, and nearly-black ℓ2 minimaxity up to a constant.
- One-group model: The one-group framework also includes three-parameter beta, normal-exponential-gamma, generalized double Pareto, generalized shrinkage, and Dirichlet-Laplace priors.These examples reinforce the section’s view that one-group models combine promising theoretical properties with computational implementation advantages.
3 The horseshoe+ estimator
The horseshoe+ estimator extends the horseshoe with an additional half-Cauchy local mixing level, producing a distinct Jacobian on the shrinkage scale. This extra U-shaped factor concentrates posterior mass near no-shrinkage and strong-shrinkage regions, strengthening ultra-sparse signal detection.
- Hierarchical construction: Horseshoe+ adds a half-Cauchy mixing variable η_i, making λ_i conditionally independent given η_i and the global parameter τ.Integrating over η_i produces a modified marginal density for the local scale, including an additional log(λ_i/τ) term.
- Shrinkage-scale behavior: The extra Jacobian term creates another horseshoe U-shaped factor that pushes posterior mass toward κ_i = 0 and κ_i = 1.These endpoints represent the regions of most interest for separating signals from noise terms on the shrinkage scale.
- Distinctive property: This additional Jacobian-based shrinkage is unique to horseshoe+ among the univariate shrinkage priors discussed.The cited comparison includes priors expressible through slowly varying functions, while horseshoe+ has an additional shrinkage effect absent from that class.
- Distinctive property: The horseshoe+ prior is the only heavy-tailed Gaussian scale-mixture shrinkage prior described as having p(κ_i) → ∞ as κ_i → 0.Its slowly varying component, together with the shrinkage-scale representation, yields extra shrinkage near the signal region.
4 Theoretical properties of the horseshoe+ estimator
The horseshoe+ prior has infinite mass near zero and sharper posterior concentration properties than the horseshoe, yielding improved convergence and asymptotically lower estimation error. Its induced testing rule attains Bayes-oracle performance up to a multiplicative constant under sparsity, with sharper type-I error bounds than the horseshoe.
- Marginal prior behavior: The horseshoe+ marginal prior density is unbounded at the origin, concentrating substantial prior mass near zero.Theorem 4.1 establishes lim |θ|→0 pHS+(θ) = ∞; comparisons indicate that horseshoe+ and horseshoe are unbounded near zero.
- Posterior concentration: The horseshoe+ posterior shrinkage weight concentrates near one as τ →0 and near zero as |yi| →∞, supporting strong null shrinkage and signal retention.These concentration properties are established through posterior inequalities for κi and hold under the stated fixed-parameter conditions.
- Testing optimality: The horseshoe+ posterior bound PHS+(κi < ϵ|yi, τ) = O(τ2) is sharper than the horseshoe bound PHS(κi < ϵ|yi, τ) = O(τ).The sharper bound does not change the asymptotic order of total Bayes risk but has implications such as lower false positives.
- Testing optimality: The horseshoe+ decision rule attains the Bayes oracle up to a multiplicative constant when the global shrinkage parameter satisfies τ = O(µ).The result is established under Assumption 1 and the theorem’s specified conditions on η and δ.
- Estimation risk: The horseshoe+ estimator has asymptotically lower posterior mean squared error than the horseshoe estimator for large |yi|.The improvement is attributed to the extra (log |yi|) factor in the marginal, arising from the extra (log λi) term in the prior mixing density.
5 Numerical examples
Numerical experiments find that horseshoe+ generally achieves the lowest estimation error, shrinks noise terms more effectively than horseshoe, and yields misclassification probabilities close to the Bayes oracle across a broad signal range.
- Estimation: The horseshoe+ prior has the lowest average SSE in all but two simulation settings, where the horseshoe prior performs best.Results use n = 200, 100 replicates, and average SSE about the posterior median.
- Estimation: The half-Cauchy prior τ ∼ C+(0, 1/n) outperforms τ ∼ Unif[0, 1] for both horseshoe and horseshoe+.Its greater mass near zero helps τ adapt to the data’s sparsity level.
- Estimation: Under τ ∼ Unif[0, 1], horseshoe+ shrinks noise terms toward zero more effectively than horseshoe.The comparison uses posterior means and middle 95% posterior credible intervals with n = 200 and 10 nonzero means equal to 7.
- Multiple testing: The horseshoe+ multiple-testing rule has misclassification probability very close to the Bayes oracle over a wide range of µ values, departing somewhat above 0.2.Misclassification probability is used because it equals Bayes risk under 0-1 additive loss in the two-groups model.
6 Application on a prostate cancer data set
The horseshoe+ prior was applied to prostate cancer gene-expression data to identify genes differentially expressed between controls and patients. Compared with the horseshoe prior, horseshoe+ estimates were closer to observed test statistics for 9 of the top 10 genes.
- Data and objective: The prostate cancer data contain expression values for 6,033 genes from 102 subjects: 50 normal controls and 52 cancer patients.The analysis seeks genes differentially expressed between controls and cancer patients.
- Data and objective: The data are analyzed as a high-dimensional normal means problem with independent Gaussian observations and simultaneous tests of H0i: θi = 0.The model assigns each observation yi a mean θi with Gaussian noise.
- Effect-size comparison: For the top 10 genes selected by Efron, the horseshoe and horseshoe+ priors were applied to all 6,033 test statistics to compare effect-size estimates.Both models used Gibbs samplers with 15,000 draws and a 3,000-draw burn-in period.
- Effect-size comparison: 9 out of the top 10 genes had horseshoe+ estimates closer to the observed test statistics than the corresponding horseshoe estimates.The authors attribute this improvement to the benefits of a heavier tail.
7 Discussion
The discussion presents horseshoe+ as an improved default shrinkage estimator for nearly-black signals, attributing its benefits to heavier tails and greater mass near zero. It also reports sharper asymptotic testing and estimation results while identifying conjectures and extensions for future work.
- Contributions: Horseshoe+ improves ultra-sparse signal separation and sparsity handling through heavier tails and greater prior mass near zero.The discussion characterizes it as a natural extension of the horseshoe prior and a default Bayesian shrinkage estimator.
- Theoretical results: Lower MSE and Bayes-oracle testing up to O(1), with a sharper type-I error bound, are established relative to horseshoe.Asymptotic minimaxity in ℓ2 is not established; the authors conjecture it may hold with a sharper constant.
- Global-local shrinkage: The global shrinkage parameter is identified as central to posterior behavior and must match the non-null proportion for horseshoe optimality.This role is linked to asymptotic minimaxity in estimation and optimality of the induced testing rule.
- Future extensions: Higher-order products of independent half-Cauchy local scales could extend horseshoe+ with heavier tails and a larger spike at zero.The discussion points to known moments and densities for products of Cauchy variables as a basis for these extensions.
- Future extensions: Adaptive one-group estimators based on Cauchy products are proposed as a conjectural extension, but their theoretical advantages remain unsettled.The authors suggest collections of different product orders could provide different shrinkage levels and extend one-group benefits.
A Proofs of theorems · A.1 Proof of Theorem 4.1
The proof derives upper and lower bounds for the horseshoe+ marginal density using a change of variables and inequalities for log(ζ)/(ζ − 1). It then shows divergence near zero and compares these bounds with those of the horseshoe prior.
- A.1 Proof of Theorem 4.1: The proof begins with the marginal density pHS+(θ) of the horseshoe+ prior, obtained from Equations (3.3–3.4) with τ = 1.
- A.1 Proof of Theorem 4.1: The horseshoe+ marginal density is represented through an integral involving λ, exp(−θ^2/(2λ^2)), log(λ), and λ^2 − 1.
- A.1 Proof of Theorem 4.1: The proof applies the transformation ζ = 1/λ^2 to analyze the resulting integral.
- A.1 Proof of Theorem 4.1: For ζ > 0, the upper-bound argument uses log(ζ)/(ζ − 1) ≤ 1/√ζ.
- A.1 Proof of Theorem 4.1: For ζ > 0, the lower-bound argument uses log(ζ)/(ζ − 1) ≥ 2/(1 + ζ).
- A.1 Proof of Theorem 4.1: Combining the bounds yields a lower and upper bound for pHS+(θ), including the upper bound 1/(π^2|θ|).
- A.1 Proof of Theorem 4.1: The lower bound diverges as θ approaches zero, establishing Part (2) of the theorem.
- A.1 Proof of Theorem 4.1: Compared with horseshoe+, the horseshoe prior has sharper upper and lower bounds, although horseshoe+ bounds may be sharpened using better approximations and an infinite-product representation.
A.2 Proof of Theorem 4.2
The proof derives bounds for the posterior shrinkage parameter κ_i by controlling a Jacobian term and bounding posterior probability ratios through integrand extremes. These steps establish an upper bound for P(κ_i < ϵ | y_i, τ).
- Posterior representation: The proof writes the posterior density of κ_i given y_i in the τ scale.
- Jacobian bound: Using 1 − 1/x < log(x) < x − 1 for x > 0, it bounds the Jacobian term by an expression involving 1 − κ_i.
- Probability-ratio bound: The posterior probability P(κ_i < ϵ | y_i, τ) is bounded by the ratio P(κ_i < ϵ | y_i, τ)/P(κ_i > ϵ | y_i, τ).
- Integral bound: The final integral bound uses the extreme values of the integrand multiplied by the interval’s integration length.
A.3 Proof of Theorem 4.3
The proof of Theorem 4.3 establishes an inequality needed for the argument and uses probability concentration to bound the type-I error rate. It approximates the posterior mean by restricting an integral as τ → 0, then derives the corresponding asymptotic expression.
- Auxiliary inequality: The proof begins by establishing an inequality required for the subsequent concentration argument.This inequality is stated as | |(κ_i(τ^2 + 1) −1)| < 1/(1 − κ_i).
- Type-I error bound: The type-I error bound applies the probability concentration inequality from Theorem 4.2 after establishing the required supporting fact.The proof explicitly frames this fact as necessary for proving the type-I error-rate bound.
- Posterior-mean approximation: The proof uses two steps to show that the posterior mean is well approximated by evaluating the integral from 1/2 to 1 as τ → 0.The displayed relation states that the restricted integral captures 1 − o(1) of the posterior contribution.
- Type-I error asymptotics: The resulting asymptotic expression is then used to calculate the type-I error rate.The passage gives the calculation through an expression involving (1 + o(1)).
A.4 Proof of Theorem 4.4 … A.7 Proof of Theorem 4.7
The appendices prove concentration and testing results through posterior shrinkage bounds, derive horseshoe+ marginal-prior asymptotics with Meijer-G identities, and compare posterior mean squared errors with the horseshoe for large observations.
- A.4 Proof of Theorem 4.4: Theorem 4.4 is proved by bounding the horseshoe+ posterior tail probability of κ_i conditional on τ and y_i.The argument compares upper and lower probability bounds using a fraction δ ∈ (0, 1) satisfying ηδ ≤ 1/(1 + τ^2).
- A.5 Proof of Theorem 4.5: Theorem 4.5 analyzes type-II error for global-local shrinkage rules under the scaling log(1/τ^2)/ψ^2 → C.Here C is identified with the threshold appearing in the Bayes-oracle risk expression.
- A.5 Proof of Theorem 4.5: The type-II error derivation applies the concentration inequality from Theorem 4.4 and tracks τ^2 C(η, δ) in the resulting bound.The proof also states that ψ^2 → C as n → ∞, with C the constant in the Bayes-oracle risk.
- A.6 Proof of Theorem 4.6: Theorem 4.6 represents the horseshoe+ marginal prior as a convolution and evaluates it using Meijer-G convolution identities.The Meijer-G class supplies convolution and Laplace-transform identities used to calculate the implied prior.
- A.6 Proof of Theorem 4.6: Power-logarithmic series expansions of G-functions determine the horseshoe+ prior’s asymptotic behavior as θ → 0 and θ → ∞.The proof identifies dominant terms around infinity and corresponding behavior around zero, then integrates near the origin to analyze Bayes risk.
- A.7 Proof of Theorem 4.7: Theorem 4.7 compares horseshoe+ and horseshoe posterior MSEs using large-|y_i| expansions for posterior means and variances.The comparison is developed from Equations (4.7)–(4.10), with the horseshoe bias controlled by Bias_HS(θ_i|y_i) = O(1/y_i^2).