Source-linked AI summary
Size, power and false discovery rates
Bradley Efron
TL;DR
Large-scale simultaneous testing requires methods that assess both false discoveries and the ability to detect nonnull cases. This paper develops an empirical Bayes false discovery rate framework, finding that power diagnostics can reveal many nonnull cases that cannot be reported without including too many null cases.
Problem
Large-scale simultaneous inference involves thousands of hypothesis tests, creating a need to assess both Type I error control and detection of genuinely nonnull cases.
Method
The paper uses empirical Bayes false discovery rate analysis with local and tail-area statistics, empirical and theoretical or permutation null estimates, and accuracy formulas for estimated rates.
Results
Power diagnostics show that a majority of nonnull cases may be unreportable as interesting without including an unacceptably high proportion of null cases.
Takeaways & Limitations
False discovery rate analysis can support size control while also diagnosing when a large-scale study has insufficient power to identify nonnull genes.
Takeaways & Limitations
In low-powered settings, any reduced thresholded list may omit investigators’ a priori favorites, motivating presentation of the full fdr values.
Abstract
from arXiv · showhide
Modern scientific technology has provided a new class of large-scale simultaneous inference problems, with thousands of hypothesis tests to consider at the same time. Microarrays epitomize this type of technology, but similar situations arise in proteomics, spectroscopy, imaging, and social science surveys. This paper uses false discovery rate methods to carry out both size and power calculations on large-scale problems. A simple empirical Bayes approach allows the false discovery rate (fdr) analysis to proceed with a minimum of frequentist or Bayesian modeling assumptions. Closed-form accuracy formulas are derived for estimated false discovery rates, and used to compare different methodologies: local or tail-area fdr's, theoretical, permutation, or empirical null hypothesis estimates. Two microarray data sets as well as simulations are used to evaluate the methodology, the power diagnostics showing why nonnull cases might easily fail to appear on a list of ``significant'' discoveries.
1. Introduction.
Large-scale simultaneous testing motivates false discovery rate methods that combine frequentist error control with empirical Bayes analysis of size and power. The paper develops these ideas through local and tail-area rates, theoretical and empirical nulls, and microarray examples showing that many nonnull cases may remain undiscovered.
- Approach: Empirical Bayes methods provide a false discovery rate framework for size and power analysis with limited frequentist or Bayesian modeling assumptions.
- Prostate example: In the prostate study, 51 of 6033 genes met the fdr ≤0.2 threshold, supporting a list expected to contain less than 20% null cases.
- Power: The estimated nonnull histogram and local fdr curve suggest low power because most nonnull cases have large fdr values and cannot be reported without many null cases.
- Null hypotheses: The theoretical null appears suitable for the prostate data but not the HIV data, where empirical null estimation is considered because the central histogram does not match N(0,1).
- Contributions: The paper considers local and tail-area false discovery rates together with theoretical and empirical null hypotheses; all combinations are possible.
2. False discovery rates.
False discovery rate analysis models large-scale test statistics through null and nonnull components, distinguishing tail-area Fdr from local fdr. The framework supports Type I error control and power assessment, while emphasizing assumptions, threshold choices, and the risk that many nonnull cases remain undiscovered.
- Model and definitions: The analysis considers N simultaneous test statistics z_i under a two-class model with null and nonnull cases occurring with probabilities p0 and p1.The null density f0 is often taken as standard N(0,1), while f1 is assumed longer-tailed; p0 and f1 may require estimation in practice.
- Model and definitions: Local fdr is the posterior probability that a case is null given z, whereas tail-area Fdr is the posterior probability of being null given Z ≤ z.Fdr is the average of fdr(Z) over Z ≤ z and is typically smaller when fdr decreases as |z| increases.
- Estimation and power: Empirical Bayes estimation uses a flexible parametric mixture-density model, with local fdr calculations supporting power diagnostics in addition to Type I error control.The paper notes moderate estimation-variability cost for estimated local fdr and uses these diagnostics to evaluate whether nonnull cases can be identified.
- Estimation and power: Thresholding can omit investigators’ a priori favorites, so reporting the full set of fdr(z_i) values is especially informative in low-powered studies.The prostate example illustrates that many nonnull cases may not enter a reduced list without including an unacceptably high proportion of null cases.
- Fdr interpretation: For the prostate data, 28 genes with z_i ≥ 3.3 imply an expected 2.71 null genes, or about one tenth of the selected cases.This interpretation requires exchangeability of the selected cases, but not independence or identical null densities, provided their average density behaves like f0.
3. Power diagnostics.
The paper develops fdr-based diagnostics for estimating power: how likely nonnull genes are to appear among low-fdr discoveries. These diagnostics also support projections of how increased sample size might change detection.
- Power diagnostics: Power diagnostics compare the estimated nonnull density with estimated fdr, using the expected fdr among nonnull cases as a simple summary statistic.The nonnull histogram provides an estimate of what would be observed for nonnull genes alone; the expectation is evaluated by numerical integration.
- Power diagnostics: An estimated nonnull expectation of fdr around 0.20 suggests good power because a typical nonnull gene is likely to appear on an interesting-candidate list.The estimate is formed from smoothed nonnull bin counts weighted by fdr values.
- Power diagnostics: The prostate study had estimated nonnull fdr expectation 0.68, while the HIV study had 0.47, indicating poor power in both examples.The prostate data’s nonnull cases mostly have large fdr values and therefore cannot be listed without also including a large percentage of null cases.
- Power diagnostics: In the simulation, 64% of nonnull cases had fdr below 0.2, compared with only 11% of prostate nonnull cases.The nonnull fdr distribution is examined through its estimated cumulative distribution, not only through its expectation.
- Sample-size projections: The diagnostics project larger-sample performance by transforming estimated nonnull counts while leaving null counts fixed, then recalculating the fdr expectation.The approach imagines c independent replicates per gene and moves nonnull counts according to the transformed mean and variance.
- Sample-size projections: Doubling the HIV study reduced estimated nonnull fdr expectation from 0.45 to 0.23, whereas doubling the prostate study produced less dramatic improvement.The reported expansion calculations use a cruder transformation with d = 1; the more detailed transformation tends to underestimate the reduction but made little difference here.
4. Empirical null estimation.
The paper estimates empirical null distributions when theoretical or permutation nulls are unreliable, using central matching and maximum-likelihood approaches. These estimates improve inferential validity but introduce accuracy and stability trade-offs.
- Motivation and null-model departures: Permutation nulls can remain close to the theoretical distribution and fail to explain the HIV data’s narrow central peak.For HIV, the permutation estimate was approximately N(0,0.992), which did not account for the observed central narrowing.
- Motivation and null-model departures: Empirical null estimation addresses distortions from unobserved covariates, array correlation, and gene correlation that can invalidate a theoretical null.The HIV data show a narrow central histogram inconsistent with N(0,1), while gene dependence need not be modeled for consistent false discovery-rate estimates.
- Consequences and trade-offs: Using an inappropriate theoretical null can undermine inferential validity, whereas empirical-null estimation avoids distortions but increases estimated false-discovery-rate variability.In the HIV data, the theoretical null left only 20 of 151 genes below the empirical-null fdr threshold of 0.20, excluding all genes with negative z-values.
- Empirical-null estimation: Maximum-likelihood fitting provides an alternative empirical-null estimate, while Figure 5 illustrates central matching for the HIV data.The fitting procedures estimate the null mean, variance, and proportion from histogram-based mixture-density estimates.
- Consequences and trade-offs: MLE fitting is generally more stable, whereas central matching is nearly unbiased for δ0 and σ0 when p0 exceeds 0.9 but can be highly variable and sensitive to discretization range.In 100 simulations, estimated standard deviations had coefficient of variation about 10% for central matching and 3% for MLE fitting.
5. Influence and accuracy.
The paper derives influence-function and delta-method formulas to quantify false discovery rate estimation accuracy across local versus tail-area and theoretical versus empirical null methods. Simulations show that accuracy depends strongly on null specification, sample size, and the estimation method.
- Accuracy framework: Closed-form influence functions support accuracy formulas for all four combinations of local or tail-area and theoretical or empirical null false discovery rates.The formulas are used to compare standard errors and estimation behavior across methodologies.
- Estimation methods: Locfdr fits a mixture density to binned z-value counts by maximum likelihood, while central matching estimates the null component from central bins.The fitted density uses a parametric family, and central matching applies least squares over a central subset.
- Assumptions and interpretation: The covariance formula is derived under independent z values, with the paper noting an application to correlated values and emphasizing that false discovery rates depend only on the order statistic.The binned count vector approximates the z-value order statistic as bin width approaches zero.
- Theoretical null: In the difficult range 2.5 ≤z ≤3.5, fdr(z) declines from 0.38 to 0.03, while local-fdr standard errors are about one third larger than tail-area Fdr errors under the theoretical null.Both methods produce stable estimates there; a 10% coefficient of variation can correspond to an estimate of 0.20±0.02.
- Empirical null: Empirical-null estimation is less accurate: a 25% coefficient of variation can yield fdr estimates such as 0.20 ± 0.05.Increasing N by factor c decreases standard errors by roughly √c, and reducing the density-estimation degrees of freedom from 7 to 5 decreased standard errors by about one third.
- Null specification: Theoretical or permutation nulls can be misleading when the null distribution is incorrect, making empirical-null estimation necessary despite reduced accuracy.For one situation, empirical-null estimates were much closer to the true fdr curve, and MLE fitting had smaller estimated standard deviation in the illustrated comparison.
6. Summary.
Large-scale simultaneous testing requires methods that address both false discoveries and the ability to detect nonnull cases. The paper shows how false discovery rate methods and empirical-null estimation support these goals, while power diagnostics reveal that many nonnull cases may remain unreportable.
- Scope: False discovery rate methods support both size and power calculations for simultaneous inference problems involving thousands of hypothesis tests.The approach brings empirical Bayes ideas to large-scale testing.
- False discovery rate methods: Local fdr statistics are better suited for Bayesian interpretation, whereas tail-area Fdr statistics were introduced for frequentist simultaneous testing.The paper analyzes both types of statistics.
- Power: Power diagnostics may show that a majority of nonnull cases cannot be reported as interesting without including an unacceptably high proportion of null cases.This illustrates why a list of significant discoveries can omit many cases that the study aims to detect.
- Empirical null: When theoretical or permutation nulls are incorrect, empirical-null estimation is presented as necessary, although it decreases the accuracy of both local and tail-area false discovery rate methods.The paper presents two empirical-null estimation methods, derives accuracy formulas, and provides the locfdr R software.