Source-linked AI summary

Optimal False Discovery Rate Control for Large Scale Multiple Testing with Auxiliary Information

Hongyuan Cao, Jun Chen, Xianyang Zhang

arXiv:2103.15311v1stat.ME

TL;DR

Large-scale multiple testing needs methods that use auxiliary information while controlling false discoveries when hypotheses have different prior null probabilities. The paper develops a shape-constrained two-group mixture and Lfdr framework with an optimal rejection rule, adaptive EM estimation, and PAVA for ordered information. It establishes asymptotic FDR control and reports finite-sample advantages over several competitors, while warning that weak auxiliary information can reduce power for strong and sparse signals.

  • Problem

    Multiple-testing procedures often assume exchangeable hypotheses, despite auxiliary information indicating differences in statistical power and prior null probability.

  • Method

    The paper estimates hypothesis-specific null probabilities and the alternative p-value density under shape constraints, then applies an optimal Lfdr-based rejection rule.

  • Results

    The method provides asymptotic FDR control and outperforms several competitors in finite samples.

  • Takeaways & Limitations

    Auxiliary information can be incorporated through an adaptive Lfdr procedure for large-scale multiple testing with ordered structure.

  • Takeaways & Limitations

    When auxiliary information is weak, the procedure can be less powerful than BH/ST, with greater loss under strong and sparse signals.

Abstract

from arXiv · show

Large-scale multiple testing is a fundamental problem in high dimensional statistical inference. It is increasingly common that various types of auxiliary information, reflecting the structural relationship among the hypotheses, are available. Exploiting such auxiliary information can boost statistical power. To this end, we propose a framework based on a two-group mixture model with varying probabilities of being null for different hypotheses a priori, where a shape-constrained relationship is imposed between the auxiliary information and the prior probabilities of being null. An optimal rejection rule is designed to maximize the expected number of true positives when average false discovery rate is controlled. Focusing on the ordered structure, we develop a robust EM algorithm to estimate the prior probabilities of being null and the distribution of $p$-values under the alternative hypothesis simultaneously. We show that the proposed method has better power than state-of-the-art competitors while controlling the false discovery rate, both empirically and theoretically. Extensive simulations demonstrate the advantage of the proposed method. Datasets from genome-wide association studies are used to illustrate the new methodology.

1 Introduction

The paper addresses multiple testing when hypotheses are not exchangeable and auxiliary information can indicate statistical power or prior null probability. It proposes an Lfdr-based framework with varying null probabilities, an optimal rejection rule, and shape-constrained estimation for ordered hypotheses.

  • Auxiliary information such as gene read counts, species relatedness, and genetic-variant features can reflect statistical power or prior null probability in multiple testing.
  • Existing auxiliary-information methods adapt thresholds or weights across hypotheses to control overall FDR, including weighted-BH, group-BH, IHW, masked-p-value, and Lfdr-based procedures.
  • The paper develops an Lfdr procedure that adaptively estimates hypothesis-specific prior null probabilities and is built on an optimal rejection rule.
  • Under correct two-group mixture specification, the proposed Lfdr procedure is expected to be more powerful than weighted-BH.
  • Theoretical results establish asymptotic FDR control and consistency results, while the method is designed to scale to large genomic datasets.
  • For ordered auxiliary information, a new EM-type algorithm jointly estimates prior null probabilities and the alternative p-value density, using PAVA under a monotone density constraint.

2 Covariate-adjusted multiple testing

The paper formulates covariate-adjusted multiple testing through a two-group mixture model with hypothesis-specific null probabilities. Its oracle rule maximizes expected true positives at a fixed marginal FDR, and its estimated step-up version has asymptotic FDR control under consistency conditions.

  • The two-group model allows the null probability π0i to vary across hypotheses while combining auxiliary information with p-value data.
  • The local FDR is the posterior probability that a hypothesis is null given its p-value, combining the hypothesis-specific prior probability with the p-value densities.
  • The oracle rejection rule maximizes expected true positives for a specified marginal FDR through cutoffs characterized by local-FDR thresholding.
  • Under a monotone likelihood-ratio assumption, local FDR is monotone in the p-value, reducing the rejection rule to thresholding p-values.
  • 2.2 Asymptotic FDR control: The oracle step-up procedure asymptotically controls FDR under conditions ensuring convergence of relevant empirical distributions and existence of a critical value.
  • 2.2 Asymptotic FDR control: The adaptive step-up procedure retains asymptotic FDR control when the estimated local-FDR quantities satisfy the stated consistency conditions.

3 Estimating the unknowns

The estimation procedure uses monotonicity constraints to estimate hypothesis-specific null probabilities and, when needed, the unknown alternative density. EM updates combined with PAVA provide a computationally efficient approach, with consistency results under shape and identifiability conditions.

  • 3.1 The density function f1(·) is known: Estimating m hypothesis-specific null probabilities requires additional constraints because unrestricted estimation is computationally prohibitive.
  • 3.1 The density function f1(·) is known: Ordered auxiliary information induces a ranked hypothesis list, and a monotone constraint makes the null-probability estimation problem solvable.
  • 3.1 The density function f1(·) is known: The proposed algorithm combines EM for the two-group mixture with PAVA to estimate ordered null probabilities and can be implemented without a tuning parameter.
  • 3.3 Asymptotic convergence and verification of Condition (C3): The pointwise estimator satisfies |π̂0,i0 − π0,i0| = Op(m^-1/3) under the theorem’s local regularity conditions.
  • 3.2 The density function f1(·) is unknown: When the alternative density is unknown, the framework permits parametric or nonparametric updates within a prespecified density class, including decreasing densities.
  • 3.3 Asymptotic convergence and verification of Condition (C3): Under the stated shape, identifiability, and regularity conditions, the estimated null probabilities and alternative density have asymptotic convergence guarantees.

4 A general rejection rule

The section develops a general rejection rule using estimated null probabilities and a decreasing surrogate for the alternative-to-null likelihood ratio. Its level-set construction is motivated by optimal local false-discovery control and is asymptotically more powerful than accumulation tests.

  • 4.1 A general rejection rule: The proposed rejection rule uses a decreasing function h as a surrogate for the likelihood ratio f1/f0, with estimated null probabilities determining hypothesis-specific thresholds.Under symmetric f0, the transformed function h(1 − x) is used in the rule.
  • 4.1 A general rejection rule: The proposed likelihood-ratio-based choice avoids subjective accumulation functions and is motivated from a Bayesian perspective.This contrasts with accumulation tests that use functions such as ForwardStop, SeqStep, and HingeExp.
  • 4.1 A general rejection rule: Unlike strict order-based procedures, the method may reject the kth hypothesis while accepting the k−1th, allowing flexibility when ordering information is weak.The authors report that this flexibility can substantially increase power when the order information is not very strong or is weak.
  • 4.1 Asymptotic power analysis: Because optimal thresholds are level surfaces of the local false-discovery rate, the proposed procedure is asymptotically more powerful than accumulation tests.The comparison is made through the asymptotic mFDR perspective.
  • 4.1 Asymptotic power analysis: The step-up procedure maximizes expected true positives among α-level FDR rules, supporting asymptotic optimal power as the number of tests grows.The asymptotic analysis is stated under regularity conditions involving the limiting rejection and false-discovery quantities.

5 Two extensions

The section extends the framework to grouped hypotheses with between-group ordering and to settings where alternative distributions vary across hypotheses. It also discusses computational compromises and identifiability constraints for estimating heterogeneous alternative distributions.

  • 5.1 Grouped hypotheses with ordering: For d ordered groups, the method imposes 0 ≤ π1 ≤ ··· ≤ πd ≤ 1 while allowing no explicit ordering within groups.The optimization can be modified by averaging estimators within each group.
  • 5.1 Grouped hypotheses with ordering: Sign information can improve power in two-sided tests when prior knowledge indicates that alternative effects are more likely positive or negative.The argument relies on sign independence from the p-value under a symmetric null distribution.
  • 5.2 Varying alternative distributions: Allowing F1i to vary freely with i makes the two-group model unidentifiable because each alternative distribution has only one informative observation.Smooth variation or shared distributions within consecutive bins provide structural compromises.
  • 5.2 Varying alternative distributions: Estimating a separate density near every index is computationally expensive for large m, motivating binning hypotheses into K consecutive groups with shared within-bin densities.For small K, the resulting computation is described as relatively efficient.
  • 5.2 Varying alternative distributions: Unlike independent hypothesis weighting, the estimated densities here directly construct the optimal rejection rule rather than only determining stratum-specific thresholds.The distinction concerns how estimated CDFs or densities enter the rejection procedure.

6 Simulation studies

Simulations evaluate FDR control and power across auxiliary-information quality, signal density and strength, dependence, noise, sparsity, and heterogeneous null or alternative distributions. OrderShapeEM generally controls FDR and is most powerful when auxiliary information is sufficiently informative, but weak auxiliary information can make BH or ST preferable.

  • 6.2 Simulation results: Removing AdaPT’s finite-sample +1 correction can cause significant FDR inflation when signal density is low.The corrected AdaPT procedure is retained for comparisons because the uncorrected variant was not reliably error-controlling.
  • 6.2 Simulation results: All procedures control the target FDR sufficiently well across settings, with no observed FDR inflation; some methods are conservative in sparse or weak-information regimes.The simulations use a pre-specified FDR level of 0.05 with empirical 95% confidence intervals.
  • 6.2 Simulation results: OrderShapeEM is overall the most powerful method when auxiliary information is not weak, whereas weak information and sparse signals can favor BH or ST.When auxiliary information is weak, incorporating it provides little benefit and OrderShapeEM may underperform those baselines.
  • 6.2 Simulation results: Power increases with signal density and signal strength across the evaluated procedures.Competitor-specific behavior varies: SABHA benefits from strong signals, while Adaptive SeqStep performs well mainly with dense signals and moderate-to-strong auxiliary information.

7 Data Analysis

The data analysis uses paired CAD GWAS datasets as reciprocal auxiliary information for identifying associated SNPs. OrderShapeEM finds more discoveries than ST and SABHA at the same FDR levels, and its discoveries substantially overlap ST while adding SNPs supported by related cardiovascular or metabolic genes.

  • 7 Data Analysis: The analysis tests 514,178 common SNPs in CARDIoGRAM and C4D, using each dataset’s p-values as auxiliary information for the other.The two analyses separately control FDR for CARDIoGRAM and C4D outcomes.
  • 7 Data Analysis: AdaPT was excluded because it did not complete the analysis within 24 hours, while BH was omitted because ST outperformed it in preliminary comparisons.The reported method comparison therefore focuses on OrderShapeEM, ST, and SABHA.
  • 7 Data Analysis: OrderShapeEM makes significantly more discoveries than SABHA and ST at the same FDR level in the C4D analysis using CARDIoGRAM as auxiliary information.The power difference becomes larger at higher target FDR levels, and the reciprocal analysis shows similar patterns.
  • 7 Data Analysis: The overlap of associated SNPs between the two datasets indicates shared genetic architecture between the populations.This overlap motivates using one GWAS as auxiliary information for the other.

8 Summary and discussions

The proposed covariate-adjusted multiple-testing procedure combines auxiliary information with cross-hypothesis p-value information through an Lfdr-based, shape-constrained framework. It offers asymptotic guarantees and finite-sample power gains, but performance depends on informative auxiliary information and assumptions about p-value distributions and dependence.

  • Summary and discussions: The oracle procedure is optimal in maximizing expected true positives for a given value of the false discovery rate.
  • Summary and discussions: The method provides asymptotic FDR control, consistency for estimating null probabilities and the alternative density, and finite-sample improvements over several methods.The estimation procedure uses isotonic regression and is tuning-parameter free and computationally fast.
  • Summary and discussions: The procedure combines auxiliary information and information across p-values to improve efficiency over competing multiple-testing methods.Its advantage is attributed to incorporating both information sources in an optimal way.
  • Summary and discussions: The method has a competitive power advantage when signals are weak and auxiliary information is moderate or strong, but can underperform BH/ST when auxiliary information is weak.The authors recommend testing prior-order informativeness before applying the method and advise against use when estimated null probabilities show little variability.
  • Summary and discussions: The procedure can experience FDR inflation when order information affects null and alternative distributions inconsistently, although such effects may be uncommon in practice.It can also lose substantial power under varying null distributions, especially with sparse signals, because it assumes uniformly distributed null p-values.
  • Summary and discussions: The marginal method may no longer be optimal under correlations, and extending it to group, spatial, or hierarchical structure remains future work.
Loading 2103.15311v1…