Source-linked AI summary
Microarrays, Empirical Bayes and the Two-Groups Model
Bradley Efron
TL;DR
High-throughput devices create thousands of simultaneous inference problems that challenge traditional hypothesis-testing frameworks. The paper examines a simple two-groups Bayesian model alongside frequentist false discovery rates, finding a relatively assumption-light z-value approach but persistent estimation costs and low power for specific cases.
Problem
High-throughput studies require simultaneous hypothesis tests for thousands of cases, creating settings in which Bayesian considerations enter large-scale frequentist inference.
Method
The paper combines the two-groups model with false discovery rates applied to z-value histograms, including theoretical and empirical null distributions.
Results
The z-value approach is notably light on assumptions, while empirical-null estimation can substantially increase variability and power to detect specific genes may remain low.
Takeaways & Limitations
Large-scale testing requires attention to null-distribution choice, estimation efficiency, and experimental design rather than relying only on individual-case detection.
Takeaways & Limitations
Empirical-null estimation increases standard errors by factors of 2 or 3 near fdr(z) = 0.2, making estimation efficiency a serious limitation.
Abstract
from arXiv · showhide
The classic frequentist theory of hypothesis testing developed by Neyman, Pearson and Fisher has a claim to being the twentieth century's most influential piece of applied mathematics. Something new is happening in the twenty-first century: high-throughput devices, such as microarrays, routinely require simultaneous hypothesis tests for thousands of individual cases, not at all what the classical theory had in mind. In these situations empirical Bayes information begins to force itself upon frequentists and Bayesians alike. The two-groups model is a simple Bayesian construction that facilitates empirical Bayes analysis. This article concerns the interplay of Bayesian and frequentist ideas in the two-groups setting, with particular attention focused on Benjamini and Hochberg's False Discovery Rate method. Topics include the choice and meaning of the null hypothesis in large-scale testing situations, power considerations, the limitations of permutation methods, significance testing for groups of cases (such as pathways in microarray studies), correlation effects, multiple confidence intervals and Bayesian competitors to the two-groups model.
1. INTRODUCTION
High-throughput technologies have transformed hypothesis testing from small simultaneous tests into large-scale inference involving thousands of cases. The article uses the two-groups model to examine frequentist and Bayesian ideas, especially False Discovery Rates, while addressing null hypotheses, power, permutations, correlation, and related problems.
- Motivation: Modern high-throughput devices require simultaneous inference across hundreds to thousands of cases, far beyond the scale of classical testing.Examples range from 200 to 10,000 cases, with larger applications anticipated.
- Motivation: When thousands of inference problems are considered together, Bayesian considerations become relevant even within frequentist analyses.The article presents the two-groups model as a simple Bayesian framework for these settings.
- Article scope: The article focuses on the interplay of Bayesian and frequentist ideas in the two-groups setting, with particular attention to False Discovery Rates.Its scope also includes null-hypothesis choice, permutation limitations, power, group significance, correlation, confidence statements, and Bayesian alternatives.
- Examples: Four applications illustrate large-scale testing: prostate genes, high schools, proteomic peaks, and brain-imaging voxels, each represented by case-specific z-values.The examples include 6033 genes, 3748 high schools, 230 peaks, and 15,445 voxels.
- Examples: The examples share a useful two-groups structure but differ in important ways, including inappropriate theoretical nulls and correlation among observations.The education data challenge the N(0,1) null, while proteomics and imaging show possible correlation patterns.
- Approach: The paper combines idealized two-groups and empirical Bayes methods, using empirical null methods when the theoretical null does not fit.The presentation aims to remain relatively nontechnical and is not intended as a comprehensive review.
2. THE TWO-GROUPS MODEL AND FALSE DISCOVERY RATES
The two-groups model gives False Discovery Rate methods a Bayesian interpretation for large-scale testing. The section connects tail-area and local false discovery rates, explains their relation to Benjamini–Hochberg control, and emphasizes that prior probabilities and modeling assumptions affect interpretation.
- Two-groups model: The two-groups model assigns each case to a null or nonnull group with prior probabilities p0 and p1 and corresponding densities f0(z) and f1(z).The mixture density combines the two groups across the cases.
- False Discovery Rates: False Discovery Rate methods have a Bayesian rationale: Fdr(z) is the posterior probability that a case is null given that its z-value falls beyond a threshold.The model uses the mixture cdf F(z) and the null cdf F0(z) to form this tail-area probability.
- False Discovery Rates: The Benjamini–Hochberg rule controls the expected proportion of null cases among reported discoveries at no more than q under independence, with extensions to some dependence models.The rule also has a Bayesian interpretation through estimated tail-area posterior probabilities.
- False Discovery Rates: Local fdr(z) is the posterior null probability at a boundary value, whereas Fdr(z) averages local false discovery rates beyond that boundary.When local fdr decreases toward the tails, Fdr is smaller than local fdr; geometrically, local fdr is a tangent slope and Fdr a secant slope.
- Interpretation: A local fdr threshold of 0.20 corresponds to q-values between 0.05 and 0.15 for moderate γ choices and can be interpreted through conservative Bayes factors.The article presents the 0.20 threshold as a subjective guide for avoiding resource-wasting discoveries.
- Interpretation: Gene-specific prior probabilities can materially change posterior null probabilities, so a hot prospect with p0(i) = 0.50 can have fdri(zi) = 0.069 when average fdr(zi) = 0.40.This illustrates why average-prior FDR assessments may not represent every investigator-selected prospect.
3. EMPIRICAL BAYES METHODS
Empirical Bayes methods estimate large-scale testing distributions from the data, combining a two-groups model with frequentist estimation. Applied to prostate data, local false discovery rates identify candidates but reveal limited power and substantial estimation costs for empirical nulls.
- Empirical Bayes estimation: Empirical Bayes estimates the marginal density f(z) from thousands of parallel cases, allowing prior information in the two-groups model to be learned from the data.The locfdr algorithm estimates the denominator f(z) using binned z-values and Poisson GLM software.
- False discovery rates: The local false discovery rate is c fdr(z) = p0f0(z)/bf(z), with only the marginal density denominator requiring estimation when p0 and f0 are specified.The theoretical-null analysis uses f0(z) = ϕ(z), while p0 near one has little effect on Fdr(z) or fdr(z).
- Prostate example: 51 of 6033 prostate genes had c fdr(zi) ≤ 0.2, compared with 60 reported by Benjamini–Hochberg at q = 0.1.The local-fdr candidates split into 26 on the right and 25 on the left; Benjamini–Hochberg reported 28 and 32, respectively.
- Power: About 422 genes were estimated to be nonnull, but only 51 were reported because most nonnull genes had c fdr(zi) > 0.5; the prostate study was underpowered.A power diagnostic gave expected nonnull fdr = 0.68, indicating low power, and a rerun could produce a substantially different list.
- False discovery rates: In the prostate bin [3.1,3.3], 17 genes versus 2.68 expected null genes produced a smoothed estimated false discovery rate of 0.24.The calculation implies that approximately one-sixth of the 17 genes in that bin were null.
4. THE EMPIRICAL NULL DISTRIBUTION
The empirical null generalizes the theoretical N(0,1) null by estimating its mean, spread, and null proportion from central z-value behavior. This matters because microarray studies can show substantial departures from N(0,1), changing false-discovery calculations and gene discoveries.
- Motivation: Large-scale testing can make the theoretical N(0,1) null visibly inappropriate, motivating empirical null distributions estimated from the data.The BRCA and HIV histograms have central peaks wider and narrower than N(0,1), respectively.
- Estimation: Empirical null methods generalize the null to N(δ0,σ0^2) and estimate δ0 and σ0 from histogram counts near z = 0.Most z-values near zero are treated as coming from null genes.
- Estimation: The geometric method fits log f(z) near zero, using a quadratic approximation to estimate p0, δ0, and σ0.The HIV analysis uses a likelihood fit followed by coefficient matching to obtain the empirical-null parameters.
- Results: For the HIV data, the theoretical null produced the impossible estimate p0 = 1.20, whereas the analytic method estimated (δ0,σ0,p0) = (−0.120,0.787,0.956).The theoretical-null fit was described as very poor.
- Methods: The analytic and geometric empirical-null methods are implemented in locfdr, with the analytic method more stable but potentially more biased than geometric fitting.Geometric fitting is nearly unbiased for δ0 and σ0 when p0 ≥ 0.90.
- Limitation: Empirical-null estimation can increase standard errors two- to threefold near fdr(z) = 0.2, because estimating δ0 and σ0 drives the added variability.Using tail-area Fdr rather than local fdr does not reduce this variability in the cited comparison.
5. THEORETICAL, PERMUTATION AND EMPIRICAL NULL DISTRIBUTIONS
The paper shows that theoretical null assumptions can fail because of nonnormality, unobserved covariates, and correlation, while permutation methods may miss these problems. Empirical null estimation can reveal and partly correct misleading false-discovery assessments.
- Why theoretical nulls fail: The theoretical null may fail when test statistics violate independence, identical-distribution, or normality assumptions.The BRCA expression matrix shows markedly nonnormal components, and across-array dependence can reduce effective degrees of freedom.
- Permutation limitations: Permutation nulls can approximate theoretical nulls and therefore fail to detect nonnormality or unobserved covariates.For BRCA, permutation results were nearly N(0,1), while the authors state that permutations cannot recognize unobserved covariates.
- Correlation effects: Correlation among z-values can make the overall null distribution substantially wider or narrower than N(0,1), even when individual null statistics are standard normal.With α = 0.153, σ0 = 1.10 and the expected null count below z = −3 is more than twice the N(0,1) expectation.
- Empirical null consequences: In simulations, the false discovery proportion averaged 0.091 at q = 0.10 but ranged from 0.03 to 0.29 across estimated null spreads.The upper-to-lower 5% averages differed by a factor of 9, showing why a nominal q-value can be misleading in a particular dataset.
- Empirical examples: The education example produced eight times as many significant schools under the theoretical null as under an estimated N(−0.35,1.512) null.The paper attributes the discrepancy to possible unobserved covariates or within-school dependence and calls the theoretical-null analysis untrustworthy.
- Empirical checks: The theoretical null can be credible, but large-scale analyses provide enough data to check it directly rather than assume it.For the prostate data, estimated parameters (bδ0,bσ0) = (0.00,1.06) were close to (0,1).
6. A ONE-GROUP MODEL
The one-group model allows true effects to vary continuously without prespecifying a null/nonnull partition. Empirical-null estimation can nevertheless recover useful local false-discovery probabilities and closely approximate full Bayes results in a blurred setting.
- Interpretation: One-group models are useful when effects range continuously from zero or near zero to large values, making a binary null/nonnull distinction less clear.The paper contrasts this setting with two-groups testing and notes that education data might reasonably be analyzed through estimation instead.
- Model: The one-group model assigns each latent effect µi a density g(µ) without requiring a separate null and nonnull component.The density may include atoms, including at zero, but no a priori partition is imposed.
- Empirical Bayes: The empirical-null fdr curve closely estimates the full-Bayes posterior probability of being uninteresting, whereas the theoretical-null curve is far off.This comparison uses the simulated model in Figure 7, where “uninteresting” means µi < 1.5.
- Empirical-null construction: Increasing the Taylor-order choice J matches the null subdensity more accurately and pushes low-fdr z-values farther from zero.The construction approximates log f(z) with higher-order Taylor expansions.
- Bayesian interpretation: The derivatives of log f(z) at zero correspond to posterior cumulants of µ given z = 0, with the second derivative adjusted by one.The first derivative is the posterior mean, while the second relates to posterior variance through the stated lemma.
- Theoretical versus empirical null: When effect mass lies near zero rather than at an atom, the empirical null differs from the theoretical null because posterior variance increases.A concentrated atom at zero yields an empirical null close to the theoretical N(0,1) null.
7. BAYESIAN AND FREQUENTIST CONFIDENCE STATEMENTS
The paper compares frequentist false-coverage intervals with Bayesian posterior intervals after selecting discoveries. FCR control achieves its target coverage, but the resulting intervals can be much wider and less well centered than Bayesian alternatives.
- FCR procedure: The Benjamini–Yekutieli FCR procedure assigns intervals to selected effects while bounding the expected proportion of noncovering intervals by q = 0.05.In the simulation, 566 discoveries were selected using z_i ≤ −2.77.
- Simulation result: Only 17 of 566 FCR intervals missed the true effect, yielding 3% noncoverage and satisfying the stated FCR property.Fourteen of the missed cases were null cases.
- Interval width: The FCR intervals were about 2 times longer than usual individual 95% intervals and were poorly centered on the left.The intervals were approximately z_i ± 2.77 rather than z_i ± 1.96.
- Bayesian comparison: Bayesian and FCR intervals pursue the same goal of including nonnull effects with 95% probability, but Bayes intervals can represent null and nonnull uncertainty separately.At z_i = −2.77, the Bayesian description assigns fdr = 0.25 to µ_i = 0 and gives a separate 95% interval conditional on nonnull status.
- Practical limitation: Estimating empirical-Bayes posterior distributions for selected nonnull effects is more challenging than estimating fdr(z).The proposed route estimates f1(z), deconvolves it to estimate the nonnull effect distribution, and then applies Bayes’ rule.
8. IS A SET OF GENES ENRICHED?
Gene-set enrichment can improve detection power, but its result depends on which null hypothesis and permutation scheme are used. Spatially correlated imaging data offer a more tractable case, where local averaging turns enrichment into spatial smoothing.
- Motivation: In the p53 data, individual-gene Fdr identified one nonnull gene, whereas enrichment analysis identified seven or eight significant gene sets.The example contains 10,100 genes measured on 50 microarrays.
- Correlation effects: Column permutations preserve correlations among pathway genes, producing a wider null distribution than row randomization.The pathway genes have highly correlated expression levels, increasing the variance of the pathway-average statistic under column permutations.
- Competing nulls: Neither permutation choice is universally correct because column permutations can instead shrink variability for BRCA data and reverse the comparison.The paper attributes the conflict to two different null hypotheses: random gene-set selection and randomized array order.
- Empirical-Bayes limitation: The two-groups empirical-Bayes approach is difficult for gene sets because mixture and null estimation become high-dimensional and prior probabilities over sets are hard to specify.The paper identifies the curse of dimensionality and the large number of possible gene sets as central obstacles.
- Imaging example: For imaging data, averaging z-values within city-block distance 2 and selecting c fdr(¯z_i) ≤ 0.2 makes enrichment equivalent to local spatial smoothing.The method applies to 15,445 voxels and exploits their convenient three-dimensional geometry.
9. CONCLUSION
Modern high-throughput science requires statistics to accommodate massive datasets and thousands of simultaneous questions, while empirical Bayes information remains incompletely understood. The paper therefore combines flexible z-value procedures with two-groups and empirical-null ideas, while emphasizing that detecting specific genes can still have low power.
- Modern applications produce massive datasets aimed at answering thousands of questions simultaneously, challenging classical hypothesis-testing frameworks.
- Empirical Bayes information accumulates across parallel but nonidentical testing situations, but how to combine that evidence remains largely unanswered.
- Massive microarray datasets can appear statistically comforting even when power to detect interesting individual genes remains quite low.
- The paper argues for new methods such as enrichment and experimental-design theory tailored to large-scale testing situations.
- Applying the two-groups model and false discovery rates to z-value histograms minimizes modeling assumptions, especially with an empirical null that does not require independence across microarrays.