Source-linked AI summary
Hypothesis Testing in High-Dimensional Regression under the Gaussian Random Design Model: Asymptotic Theory
Adel Javanmard, Andrea Montanari
TL;DR
The paper addresses the less-understood problem of testing individual coefficients in high-dimensional sparse regression, rather than only estimating or selecting them. It uses debiasing of the Lasso, proves minimax power bounds, and shows near-achievability for Gaussian designs across distinct sample-size regimes. The main guarantees are strongest for standard designs at n of order s0 log(p/s0), while general designs require a stronger rigorous regime or replica-based justification.
Problem
High-dimensional regression lacks well-understood methods for computing p-values and assessing statistical significance of individual coefficients.
Method
The paper combines Lasso debiasing with minimax testing bounds and asymptotic distributional characterizations for standard and general Gaussian designs.
Results
For standard Gaussian designs, the proposed test is nearly optimal at n of order s0 log(p/s0), with detectable signals of order σ/√n; for general designs, the distributional limit is proved when n is much larger than s0(log p)^2.
Takeaways & Limitations
The results provide a practical route to statistically significant coefficient testing near the point-estimation sample-size threshold for standard Gaussian designs.
Takeaways & Limitations
The analysis is asymptotic, and general-design results require knowing or estimating the covariance matrix Σ; proving the standard distributional limit in full generality remains challenging.
Abstract
from arXiv · showhide
We consider linear regression in the high-dimensional regime where the number of observations $n$ is smaller than the number of parameters $p$. A very successful approach in this setting uses $\ell_1$-penalized least squares (a.k.a. the Lasso) to search for a subset of $s_0< n$ parameters that best explain the data, while setting the other parameters to zero. Considerable amount of work has been devoted to characterizing the estimation and model selection problems within this approach. In this paper we consider instead the fundamental, but far less understood, question of \emph{statistical significance}. More precisely, we address the problem of computing p-values for single regression coefficients. On one hand, we develop a general upper bound on the minimax power of tests with a given significance level. On the other, we prove that this upper bound is (nearly) achievable through a practical procedure in the case of random design matrices with independent entries. Our approach is based on a debiasing of the Lasso estimator. The analysis builds on a rigorous characterization of the asymptotic distribution of the Lasso estimator and its debiased version. Our result holds for optimal sample size, i.e., when $n$ is at least on the order of $s_0 \log(p/s_0)$. We generalize our approach to random design matrices with i.i.d. Gaussian rows $x_i\sim N(0,Σ)$. In this case we prove that a similar distributional characterization (termed `standard distributional limit') holds for $n$ much larger than $s_0(\log p)^2$. Finally, we show that for optimal sample size, $n$ being at least of order $s_0 \log(p/s_0)$, the standard distributional limit for general Gaussian designs can be derived from the replica heuristics in statistical physics.
1 Introduction
The paper studies p-values and hypothesis tests for individual coefficients in high-dimensional sparse regression, where p exceeds n. It derives minimax limits and develops nearly optimal procedures for standard and general Gaussian designs under different sample-size regimes.
- Problem: The paper tests whether individual sparse-regression coefficients are nonzero, quantifying the trade-off between significance level α and power 1 −β.Without a minimum signal magnitude, nonzero coefficients can be arbitrarily close to zero and indistinguishable from noise.
- Problem: For s0-sparse coefficients, the paper seeks necessary and sufficient conditions on n, p, s0, σ, and µ for prescribed testing significance and power.The target is uniform testing over coefficient vectors with |θ0,i| > µ.
- Benchmark: In orthogonal designs, reliable testing requires |θ0,i| of order σ/√n, establishing the benchmark signal scale.For fixed α and β, the threshold is c(α, β)σ/√n.
- Standard Gaussian designs: The paper answers positively for standard Gaussian designs by combining Lasso-based debiasing with minimax power bounds.It analyzes the challenging regime where the target coefficient is at the noise level.
- Standard Gaussian designs: For standard Gaussian designs, the lower and upper signal thresholds both scale as σ/√n, up to a universal constant, when n is of order s0 log(p/s0).The proposed test is nearly optimal in its power-significance trade-off and answers the stated open question.
- General Gaussian designs: For general Gaussian covariance Σ, the method is rigorously supported when n is much larger than s0(log p)^2, while the optimal-sample-size limit is derived using replica heuristics.Conditional on the generalized distributional limit, the procedure is nearly optimal.
- Limitations: The treatment is asymptotic and, for general designs, requires knowing or estimating the covariance matrix Σ.A full study of covariance estimation is beyond the paper’s scope; later work uses larger sample sizes under weaker design assumptions.
2 Minimax formulation
The paper formulates single-coefficient inference as a minimax testing problem, balancing significance against power uniformly over sparse parameters. It derives an upper bound through binary-testing and oracle reductions, showing that Gaussian designs can attain the ideal detection scale under suitable conditions.
- 2.1 Tests with guaranteed power: A test rejects H0,i when Ti,X(y) = 1, with the design matrix X treated as part of the testing procedure's input.The procedure is represented by measurable functions indexed by the coefficient and design.
- 2.1 Tests with guaranteed power: The framework evaluates tests by significance level α and power 1 −β, corresponding respectively to false-positive and false-negative control.The minimax formulation requires these guarantees uniformly over s0-sparse parameter vectors.
- 2.1 Tests with guaranteed power: The minimax power 1 −βopt_i(α; µ) is the best worst-case power at significance level α against alternatives satisfying |θi| ≥µ.For standard Gaussian designs, exchangeability makes the criterion independent of the tested index i.
- 2.2 Upper bound on the minimax power: The upper-bound function G(α, u) is continuous and increasing in both its arguments, with G(α, 0) = α and limit 1 as u grows.A randomized test establishes the baseline lower bound 1 −βopt_i(α; µ) ≥ α.
- 2.2 Upper bound on the minimax power: For general Gaussian covariance, optimal power is bounded by scalar-Gaussian mean testing with effective variance σ²_eff ≈ σ²/[Σi|S(n−s0)].The bound uses covariance residualization through Σi|S and chi-squared concentration.
- 2.2 Upper bound on the minimax power: For standard Gaussian designs, achieving target pair (α, β) requires µ ≥ cσ/√n, establishing the ideal detection scale for sparse-coordinate testing.The paper compares this scale with prior procedures requiring an additional s0 log p/n term.
- 2.2 Upper bound on the minimax power: The upper bound uses an oracle that knows the support excluding coordinate i, yet the paper states that this apparently strong bound is asymptotically tight.The oracle argument reduces arbitrary testing procedures to a more informed benchmark.
- 2.3 Proof outline: The proof reduces minimax testing to binary testing by placing distributions Q0 and Q1 on sparse null and alternative parameter classes, then reducing the resulting problem to scalar regression.Lemmas 2.6 and 2.7 provide the reduction, using projections onto and orthogonal to selected design columns.
3 Hypothesis testing for standard Gaussian designs
For standard Gaussian designs, SDL-test debiases the Lasso to obtain asymptotically valid p-values and nearly optimal power. Its guarantees rely on the asymptotic distribution of the debiased estimator and apply in the high-dimensional regime under stated sparsity and scaling conditions.
- SDL-test is a practical procedure for testing individual coefficients under the standard Gaussian design model.The procedure is described in Table 1 and is designed for high-dimensional regression.
- Asymptotic characterization: The debiased estimator bθu is asymptotically unbiased and normally distributed, with variance τ^2 consistently estimable from the residual vector.Under the null, its coordinates are asymptotically Gaussian with mean 0 and variance τ^2, motivating a two-sided rejection threshold.
- Asymptotic characterization: The debiasing step modifies the Lasso estimator in the direction of increasing its ℓ1 norm to eliminate its bias.This construction has a geometric interpretation tied to the Lasso subgradient.
- Asymptotic guarantees: Theorem 3.2 establishes asymptotic control of type I errors for the resulting p-values along converging sequences of standard Gaussian instances.The analysis uses the support of θ0 and an asymptotic distributional characterization of the Lasso estimator.
- Asymptotic guarantees: For n ≥ 2s0 log(p/s0)(1 + O(s0/p)), the required signal magnitude is O(σ/√n), matching the fundamental limit for testing.Theorem 3.3 gives the power result, and the paper states that this matches the lower bound established earlier.
- Numerical experiments: SDL-test achieves higher statistical power than ridge-based regression and LDPE at comparable type I error levels, while those broader-design methods are overly conservative.Experiments report agreement with the asymptotic prediction, and the comparison in Figure 3 holds empirical type I error fixed across methods.
4 Hypothesis testing for nonstandard Gaussian designs
The paper extends SDL-test hypothesis testing to Gaussian designs with general covariance by using a standard distributional limit, rigorously established at larger sample sizes and supported by replica heuristics in a broader regime.
- 4.1 Hypothesis testing procedure: The generalized SDL-test applies to random designs with independent Gaussian rows x_i ∼ N(0, Σ), initially assuming known covariance Σ.The procedure is defined through a debiased Lasso estimator and p-values based on an asymptotically standard-normal statistic.
- 4.2 Asymptotic analysis: A standard distributional limit requires the debiased estimator’s empirical coordinates and normalized residuals to converge to specified limiting distributions.The definition uses parameters τ and d, with normalized residuals converging weakly to N(0, 1).
- 4.2 Asymptotic analysis: The generalized SDL-test controls type I errors and characterizes power whenever the standard distributional limit holds.The power theorem applies under additional sparsity and signal-size assumptions, and its result holds for any regularization parameter λ.
- 4.3 Gaussian limit for n ≫ s0(log p)^2: When n asymptotically dominates s0(log p)^2, the standard distributional limit can be established rigorously under bounded-eigenvalue and regularity assumptions on Σ.The resulting limit has d = (1 − ||bθ(λ)||0/n)^−1 and τ = σ0.
- 4.4 Gaussian limit via replica heuristics: Replica heuristics indicate that the standard distributional limit extends to a large class of non-diagonal covariance structures, including suitable block-diagonal matrices.The paper emphasizes that this evidence is not fully rigorous, while numerical experiments support the usefulness of the framework.
- 4.5 Covariance estimation: A covariance-free alternative uses estimated bounds on Σ, while numerical experiments show SDL-test type I errors closely match the chosen significance level.The main SDL-test instead requires covariance estimation, and the experiments report robustness to covariance-estimation errors.
5 Discussion
The discussion contrasts the paper’s debiasing approach with related methods and explains why the scaling factor d is essential for Gaussian behavior. It also summarizes simulation evidence and the method’s scope relative to competing sample-size and covariance assumptions.
- Numerical comparisons: The numerical experiments compare SDL-test with ridge-based regression and LDPE across standard and nonstandard Gaussian-design settings.The discussion includes histograms, kurtosis comparisons, and simulations of type I error and power.
- Comparison with other debiasing methods: The paper’s rigorous results are fully established only for the special case Σ = I, while general Gaussian designs require stronger assumptions and covariance knowledge.The general-design extension relies on the standard distributional limit assumption and knowledge or estimation of Σ.
- Comparison with other debiasing methods: SDL-test achieves similar power and confidence intervals with optimal sample size n = O(s0 log(p/s0)).Related methods discussed require sample sizes asymptotically larger than (s0 log p)^2.
- Role of the factor d: Without the scaling factor d, the normalized estimator’s empirical kurtosis does not converge to zero, whereas the debiased version is asymptotically Gaussian.The effect of d is strongest when n is close to s0 and becomes less noticeable when n = 30s0.
6 Real data application
The paper applies SDL-test to the UCI Communities and Crimes dataset after preprocessing the attributes and constructing a high-dimensional subsampling experiment. The application evaluates type I error, power, and selected relevant features against ridge-based regression.
- Dataset and preprocessing: The Communities and Crimes dataset contains 1,994 communities, 122 predictive attributes initially, and a violent-crime response variable.After removing 16 attributes, the analysis uses p = 106 predictors.
- Dataset and preprocessing: The preprocessing replaces missing values by attribute means, removes linearly dependent attributes, and normalizes columns to mean zero and ℓ2 norm √ntot.These steps produce Xtot ∈ R^1994×106.
- Hypothesis-testing evaluation: The study defines active variables using the full-data least-squares estimator, treating coefficients with magnitude larger than 0.04 as active.The remaining coefficients are treated as inactive when computing type I errors and power.
- Hypothesis-testing evaluation: The experiment uses random subsamples of n = 84 communities and compares SDL-test with ridge-based regression over 20 realizations.The evaluation reports type I error and statistical power at α = 0.01, 0.025, and 0.05.
- Results: The reported results include relevant-feature comparisons between SDL-test and ridge-based regression, with false-positive predictions identified in the feature table.The application also examines normalized histograms for active and inactive coordinates.
7 Proofs
The proofs establish minimax testing bounds through reductions to binary Gaussian hypothesis testing and derive distributional and testing results from standard distributional limits. The arguments also verify the required design properties and scaling behavior for the debiased estimator.
- Minimax testing bound: The minimax power upper bound is proved by reducing coefficient testing to binary hypothesis testing under carefully chosen Gaussian priors.The proof uses the most powerful likelihood-ratio test for a Gaussian mean problem with known covariance.
- Minimax testing bound: The power of the constructed Gaussian test converges to the oracle power as the auxiliary variance parameter a tends to infinity.Taking a sufficiently large establishes the claimed bound.
- Distributional limits: Under a standard distributional limit, the empirical distribution of the true and debiased coefficients converges weakly to a Gaussian-noise representation.This convergence supports asymptotic evaluations of testing errors and power.
- Testing guarantees: For standard Gaussian designs, exchangeability of the columns yields asymptotic type I error equal to the target significance level for active coordinates.The proof takes expectations after establishing the relevant empirical convergence.
- Technical ingredients: The proof verifies restricted eigenvalue conditions and controls the debiasing factor by showing that ∥bθ∥0/n tends to zero.Consequently, d can be confined to a bounded interval for sufficiently large dimensions.
A Effective noise variance τ 2
This appendix explains the effective noise variance τ^2 of the debiased estimator. It characterizes τ0 through equations involving the scalar denoising model and states existence and uniqueness on a specified interval.
- Effective noise variance: The unbiased estimator bθu is asymptotically a noisy version of θ0 with noise variance τ0^2.The paper provides an explicit formula for τ0 in the scalar characterization.
- Scalar characterization: The scalar characterization uses the soft-thresholding function with threshold a.The function is piecewise defined as x − a for x > a, zero for −a ≤ x ≤ a, and x + a otherwise.
- Scalar characterization: τ0 is obtained by solving two equations for κ and τ with κ restricted to (κmin, ∞).κmin(δ) is defined as the unique non-negative solution of a preceding equation.
- Scalar characterization: Existence and uniqueness of τ0 are established in the cited prior result.The appendix uses that result rather than reproving existence and uniqueness.
B Tunned regularization parameter λ
The appendix tunes the regularization parameter λ by optimizing a soft-thresholding risk over ε-sparse distributions, then mapping the optimal threshold through the model equations.
- M(ε, κ) is the minimax risk of soft thresholding at threshold κ over the family of ε-sparse distributions.The function is evaluated on the worst-case ε-sparse distribution.
- M(ε, κ) = ε(1 + κ^2) + (1 − ε)[2(1 + κ^2)Φ(−κ) − 2κφ(κ)].
- κ*(ε) is the minimax-optimal threshold over ε-sparse distributions.
- The theorem’s λ is obtained by solving Eq. (103) for τ at κ = κ*(ε), then substituting κ*(ε) and τ into Eq. (104).This yields λ = λ(pΘ0, σ, ε, δ).
- In the standard Gaussian setting, Eq. (104) is equivalent to another characterization for converging sequences of instances.The normalization factor d is specified by Eq. (17).
C Statistical power of earlier approaches
The appendix compares power requirements for earlier inference procedures under deterministic or standard Gaussian designs. These comparisons show how design assumptions and correction strategies constrain detectable signal strength.
- Earlier deterministic-design procedures require a significantly larger μ/σ to control both type I and type II errors.This comparison concerns restricted-eigenvalue designs.
- LDPE constructs confidence intervals through a low-dimensional projection estimator, but its rejection condition is constrained by lower bounds on τj and ε′n.For standard Gaussian designs, η* ≥ √log p and the projection norm is bounded with high probability.
- The corrected ridge estimate removes the projection-bias term PXθ0 − θ0 in addition to the regularization bias.The correction is motivated by decomposing ridge-estimator bias into these two components.
- Even assuming ℓ1-consistency of the corrected estimate, rejecting H0,j with non-negligible probability requires an additional signal-strength condition.
- For standard Gaussian designs, a necessary condition for rejecting H0,j with non-negligible probability follows from the behavior of projected design coordinates and Ωjj.The relevant quantities concentrate or satisfy high-probability bounds.
D Replica method calculation
The replica calculation derives the standard distributional limit by analyzing a generalized separable-penalty estimator through a concentrated partition function and saddle-point equations. The derivation then verifies the claimed limit for the debiased estimator.
- The analysis replaces the Lasso with regularized least squares using an arbitrary convex separable penalty J(θ).Ridge and the Lasso are included as special cases, and the corresponding factor d is defined through a saddle-point equation.
- For a fixed sequence of Gaussian-design instances, the calculation chooses d̃ = 1/ζ and compares the resulting derivative with the target characterization.This comparison establishes the standard distributional limit, with τ² and d matching the claimed formulas.
- The replica calculation studies a partition function whose normalized log is assumed to have almost-sure limits as p and β grow, with the limits exchangeable.The generating function is built from the debiased estimator and a convex dual representation.
- The derivation assumes concentration of the normalized log-partition function and permits differentiation through the limiting operation under a convexity condition.The stated sufficient example is strong convexity of the cost at s = 0.
- The debiased estimator is written as bθu = bθ + (d̃/n)Σ^-1Xᵀ(y − Xbθ), with bθ minimizing the regularized least-squares objective.The replica calculation uses duality and obtains the joint empirical distribution of the debiased estimate, true parameter, and diagonal precision term.
- Replica symmetry imposes permutation invariance on the saddle point, after which saddle-point equations determine the limiting parameters.The ansatz is motivated by invariance of the action, while correctness is expected for the log-concave partition function.
E Simulation results
The simulations compare SDL-test with ridge-based regression, LDPE, and theoretical power bounds across standard and nonstandard Gaussian designs. Results are summarized at two significance levels using repeated realizations.
- SDL-test, ridge-based regression, and LDPE are evaluated for type I error and statistical power over 10 realizations per configuration.The identity-covariance experiments are compared with the asymptotic bound from Theorem 3.3.
- Circulant-covariance experiments compare the same procedures with the lower bound for SDL-test from Theorem 4.4.
- Each table reports means and standard deviations of type I errors and powers across 10 realizations.Configurations are encoded by (p, n, s0, μ), such as (1000, 600, 50, 0.1).
- At α = 0.05, Table 8 covers standard Gaussian design and Table 10 covers the described circulant covariance design.
- At α = 0.025, Table 9 covers standard Gaussian design and Table 11 covers the described circulant covariance design.
F Alternative hypothesis testing procedure
The alternative hypothesis-testing procedure avoids directly estimating the full covariance matrix by using bounds on its off-diagonal entries. It connects to an earlier testing procedure while deriving p-values from a standard distributional limit for debiased Lasso estimates.
- Procedure: The alternative procedure requires bounds on Σ that can be estimated from the data, rather than an estimate of the full covariance matrix.It is presented as an alternative to SDL-test and is not the section’s main focus.
- Distributional basis: The standard distributional limit motivates treating transformed debiased Lasso coordinates as asymptotically Gaussian, with a normalized variable independent of the signal distribution.The paper defines the transformed coordinate using Σ, its diagonal entries, and the asymptotic scale τ.
- P-value construction: The testing construction seeks constants ∆i that make the test statistic stochastically bounded by a standard normal variable, enabling two-sided p-values.The p-value construction then yields control of type I errors.
- Covariance control: The procedure only needs bounds on maxj≠i |Σij|, which can be obtained from the empirical covariance matrix bΣ = (1/n)X^TX.The argument uses an entrywise empirical-covariance bound under Gaussian design.
- Covariance control: The resulting covariance bound follows by comparing the maximum true off-diagonal covariance with the corresponding empirical maximum and the maximum estimation error.The comparison is expressed through the inequality maxj≠i |Σi,j| − maxj≠i |bΣi,j| ≤ maxj≠i |bΣi,j − Σi,j|.
G Proof of Lemma 7.1
The proof controls Gaussian-design bilinear forms uniformly over sparse supports using sub-exponential concentration, ε-nets, and union bounds. It establishes the stated high-probability bound for all relevant pairs of subsets.
- Gaussian concentration: For sparse unit vectors supported on A and B, the proof studies centered variables ξi = ⟨u,Kxi⟩⟨xi,v⟩−⟨u,v⟩.Here K = Σ−1, and the variables are independent with mean zero.
- Gaussian concentration: The variables ξi have uniformly bounded sub-exponential norms because Gaussian transformations are sub-Gaussian and products of sub-Gaussian variables are sub-exponential.The bound depends on constants determined by cmin and cmax.
- Concentration bound: When n = ω(s0 log p), Bernstein’s inequality provides the required concentration for each fixed pair of sparse unit vectors.The probability bound is then extended from fixed vectors to sparse sets.
- ε-net argument: A 1/2-net of each sparse unit sphere has size at most 5^|A| and 5^|B|, allowing a union bound over the net vectors.The supports A and B determine spheres isometric to S^|A|−1 and S^|B|−1.
- Uniform bound: The net argument yields probability at least 1 − 5^|A|+|B|p^−c1s0 for the intermediate uniform bound.The same probability statement appears after applying the lemma and bound in the proof.
- Uniform bound: A union bound over fewer than p^2c0s0 pairs of supports produces the final high-probability statement with K depending on c0, cmin, and cmax.The conclusion applies simultaneously to all pairs with |A|, |B| ≤ c0s0.