Source-linked AI summary
Gaussian Graphical Model Estimation with False Discovery Rate Control
Weidong Liu
TL;DR
High-dimensional GGM methods rely on tuning parameters, but their relationship to false edges is difficult to derive. This paper introduces GFC, using bias-corrected residual-covariance test statistics and simultaneous thresholding, and shows asymptotic FDR/FDP control. Simulations report favorable performance, while detecting very small nonzero entries remains difficult.
Problem
Regularized GGM methods require tuning parameters, but the precise relationship between those parameters and the number of false edges is difficult to derive.
Method
GFC uses bias-corrected residual-covariance test statistics and simultaneous multiple testing for conditional dependence in high-dimensional GGM estimation.
Results
GFC controls both FDR and FDP asymptotically.
Takeaways & Limitations
The proposed procedure provides a multiple-testing alternative for evaluating high-dimensional GGM edges through asymptotic false-discovery control.
Takeaways & Limitations
The numerical results indicate that very small nonzero entries are difficult to detect.
Abstract
from arXiv · showhide
This paper studies the estimation of high dimensional Gaussian graphical model (GGM). Typically, the existing methods depend on regularization techniques. As a result, it is necessary to choose the regularized parameter. However, the precise relationship between the regularized parameter and the number of false edges in GGM estimation is unclear. Hence, it is impossible to evaluate their performance rigorously. In this paper, we propose an alternative method by a multiple testing procedure. Based on our new test statistics for conditional dependence, we propose a simultaneous testing procedure for conditional dependence in GGM. Our method can control the false discovery rate (FDR) asymptotically. The numerical performance of the proposed method shows that our method works quite well.
1 Introduction
High-dimensional GGM estimation is difficult because regularization requires tuning parameters whose relationship to false edges is unclear. The paper proposes simultaneous conditional-dependence tests with asymptotic FDR/FDP control and favorable simulated performance.
- Motivation: GGM estimation recovers the support of the precision matrix, representing conditional dependence among jointly Gaussian variables.
- Problem: Regularization methods require tuning parameters: large values can miss small-weight edges, whereas small values can generate many false edges.
- Problem: The precise relationship between tuning parameters and the number of false edges is difficult to derive, limiting rigorous performance evaluation.
- Proposed approach: The paper proposes GFC, a multiple-testing procedure for high-dimensional GGM estimation that directly tests conditional dependence.
- Proposed approach: GFC introduces bias-corrected residual-covariance test statistics that are asymptotically normal under sparsity conditions on Ω.
- Results: GFC controls FDR and FDP asymptotically, has computational cost comparable to neighborhood selection or CLIME, and performs favorably in simulations.
2 Tests on conditional dependence
The paper constructs high-dimensional conditional-dependence tests from bias-corrected residual covariance estimates and applies them simultaneously across all variable pairs. Thresholding these statistics yields asymptotic FDR and FDP control under stated conditions.
- Multiple testing: The new statistics permit simultaneous testing of (p^2 − p)/2 hypotheses, with the null limiting distribution free of unknown parameters.
- Implementation: The initial regression coefficient estimators can be obtained with Lasso or Dantzig selectors under convergence-rate and sparsity conditions.
- Multiple testing: GFC thresholds test statistics directly and estimates the null set using sparsity of Ω rather than relying on true p-values.
- Theory: Under the paper’s dependence and sparsity framework, FDR(t) converges to α and FDP(t) converges to α in probability.
3 Theoretical results
Theoretical results establish that GFC asymptotically controls false discovery proportion and false discovery rate under high-dimensional sparsity and regularity conditions.
- The procedure's FDR control is established under bounded covariance and precision-matrix diagonals with log p = o(n).These are part of condition (C1).
- GFC controls FDP and FDR at level α asymptotically.Theorem 3.1 provides the central control guarantee.
- The sparsity condition max_i Card(A_i(γ)) = O(p^ρ), ρ < 1/2, limits the number of sufficiently large precision-matrix entries per row.The paper describes this condition as mild and notes it holds under row sparsity O(√n) in a stated regime.
- The dimension p may be much larger than the sample size n because the theorem permits arbitrarily large r in p ≤ n^r.This conclusion is stated alongside the theorem's asymptotic guarantee.
- The condition for controlling FDR may be weaker than the condition for controlling FDP.Even if condition (14) is violated, FDR may still be controlled at level α.
4 Data-driven choice of ˆβi
The paper develops estimator choices for the regression coefficients required by GFC, including fixed and adaptive tuning options that produce theoretically suitable estimators.
- GFC requires estimators of β_i, and the paper uses Dantzig and Lasso estimators as available choices.Scaled-Lasso and Square-root Lasso are also identified as alternatives with similar theoretical results.
- Setting δ = 2 in the Lasso construction yields estimators satisfying conditions (3), (4), and (13).The same type of guarantee is stated for any δ > 2 in the alternative construction.
- For finite samples, the paper proposes selecting δ adaptively from the data rather than fixing it at a potentially large value.The selected estimator is β̂_i(δ̂), obtained after constructing the corresponding test statistics.
- The adaptive choice minimizes an error involving null and nonzero test-statistic contributions, subject to α ≥ τ with τ bounded away from zero.The implementation sets τ = 0.3 and discretizes the resulting integral.
- Theoretical properties of the data-driven selector δ̂ are left for future work.The paper identifies deriving those properties as important but does not provide them here.
5 Numerical results
Simulations evaluate GFC across Band, Hub, and Erdős–Rényi graphs, comparing false-discovery control and power with GFC-Dantizg, GFC-Lasso, and Glasso.
- FDR control: GFC procedures control FDR near or below α across all three graph types.GFC-Dantizg FDRs are close to α for Band and Erdős–Rényi graphs and somewhat smaller for Hub graphs.
- Tuning selection: Several tuning-index choices yield estimated FDRs well controlled at α = 0.2, and the selected index can take these values across all three graphs.This pattern is reported for GFC-Dantizg and similarly observed for GFC-Lasso.
- Power: Power increases with α; it is close to one for Hub graphs, while GFC-Lasso is more powerful than GFC-Dantizg for Band graphs.Erdős–Rényi power is non-trivial for p = 50, 100, 200 but low when p = 400.
- Power: When p = 400 in the Erdős–Rényi graph, standardized precision-matrix entries lie in (0.1275, 0.255), making nonzero edges difficult to detect.This provides the reported explanation for the low power at p = 400.
- Comparison with Glasso: For Band and Erdős–Rényi graphs, GFC power is at most 0.05 when Glasso FDR is at most 0.2, so GFC outperforms Glasso in these settings.For Hub graphs, Glasso power is close to one at small FDRs, similarly to GFC.
6 Proof
The proof establishes the paper’s asymptotic results through Gaussian approximation, concentration bounds, covariance conditions, and combinatorial decompositions of dependent index pairs.
- Assumptions: The argument assumes independent mean-zero random vectors with finite moment conditions and bounded pairwise covariance.It also imposes covariance-approximation conditions involving logarithmic powers of p.
- Threshold and FDR control: A data-dependent threshold is shown to lie in the required range with probability tending to one, after which the false-discovery ratio is controlled in probability.The proof uses q0 = Card(H0) and auxiliary lemmas to establish the required bounds.
- Regression estimation: The regression-estimation propositions verify concentration and restricted-eigenvalue conditions needed for the coefficient estimators.The restricted eigenvalue constant is lower-bounded using λmin(Σ).