Source-linked AI summary

Comment: Microarrays, Empirical Bayes and the Two-Groups Model

Carl N. Morris

arXiv:0808.0597v1stat.ME

TL;DR

The paper addresses how Bayes, frequency, and empirical Bayes reasoning can support simultaneous inference for massive datasets without being confined to exchangeable settings. It develops probabilistic models and extensions for identifying substantial random effects, connecting microarray analyses with hospital profiling. The discussion concludes that nonexchangeability, multilevel modeling, and regression toward the mean require further attention.

  • Problem

    Empirical Bayes analyses and simultaneous-inference methods need to address massive datasets and settings where exchangeability of data and random effects is restrictive.

  • Method

    The paper combines frequency, Bayesian, and empirical Bayes perspectives, extending two-groups modeling toward posterior identification of random effects exceeding a substantive threshold.

  • Results

    The paper shows how estimated mixing distributions can support threshold-based identification of large random effects and relates nonexchangeability to long-tailed school-data behavior.

  • Takeaways & Limitations

    Gene identification and hospital profiling can benefit from decision models and multilevel empirical Bayes models that account for scientifically meaningful departures and regression toward the mean.

  • Takeaways & Limitations

    The analyses rely substantially on exchangeability, while summarized school data can obscure nonexchangeability caused by differing amounts of data across rows.

Abstract

from arXiv · show

Brad Efron's paper [arXiv:0808.0572] has inspired a return to the ideas behind Bayes, frequency and empirical Bayes. The latter preferably would not be limited to exchangeable models for the data and hyperparameters. Parallels are revealed between microarray analyses and profiling of hospitals, with advances suggesting more decision modeling for gene identification also. Then good multilevel and empirical Bayes models for random effects should be sought when regression toward the mean is anticipated.

1. FREQUENCY, BAYES, EMPIRICAL BAYES AND A GENERAL MODEL

The paper frames frequency, Bayesian, and empirical Bayes perspectives within a general hierarchical model, while questioning exchangeability as a necessary condition for empirical Bayes analyses. Massive datasets create opportunities for richer models and more accurate inference.

  • Massive datasets: Massive datasets enable statisticians to develop better models that improve inference for observed data and advance future scientific discoveries.
  • Exchangeability: Brad Efron’s empirical Bayes analyses use six datasets treated as exchangeable, whereas earlier work sought shrinkage methods without that requirement.Nonexchangeability matters when estimated random effects have different variances, such as from different sample sizes.
  • General model: A general hierarchical model includes distributions for data given parameters and distributions for hyperparameters governing those parameters.The framework can extend to infinite-dimensional parameters or hyperparameters.
  • General model: Frequency and Bayesian models form endpoints of a continuum whose middle accommodates flexible empirical Bayes restrictions on distributions.Decision theory permits conditional frequency evaluations across hyperparameter values.
  • Frequency and Bayes: Bayesian reasoning can guide construction of frequency procedures, and admissible procedures nearly coincide with extended Bayes rules.

2. FDR, FDR AND EXCHANGEABILITY

The paper develops false discovery measures from probabilistic two-groups modeling of exchangeable hypothesis-testing problems. Local fdr estimates the posterior probability of the null, while Fdr averages that quantity under a threshold rule, but validity depends on restrictive exchangeability assumptions.

  • Two-groups model: The two-groups model specifies null and alternative densities, f0 and f1, for simultaneous hypothesis testing.The approach progresses through increasingly elaborate probability models for exchangeable data and repeated problems.
  • Local fdr: In the simplest exchangeable case, estimating P(H0|z) requires p0, f0, and the marginal density f(z).The marginal density satisfies f(z) = p0 ∗f0(z) + p1 ∗f1(z).
  • Limitations: The simple marginal-density approach may be valid for five of the six datasets, but not for the school data.
  • Local fdr and Fdr: Local fdr is the posterior probability of H0 given z, whereas Benjamini–Hochberg Fdr is its integral under a threshold-based decision rule.
  • Exchangeability: The repeated-testing probability model assumes exchangeability through a common p0 and common f0 and f1 distributions across problems.Conditional densities may depend on i through the random effect µi, but not otherwise on i.
  • Exchangeability: Exchangeability permits estimating Bayesian procedures from the marginal distribution of observed zi, but it can be restrictive and fails for school data with varying enrollments.

3. MULTIPLE HYPOTHESIS TESTING—LOOKING FOR LARGE RANDOM EFFECTS

The section extends two-groups and empirical-Bayes reasoning from testing µ = 0 toward identifying effects exceeding a scientifically meaningful threshold k. It connects posterior probabilities, estimated mixing distributions, and multilevel decision-making for genes and hospital performance.

  • Threshold-based identification: The posterior distribution p(µ|z) supports selecting effects with P(µ ≥ k|z), rather than testing only the point null µ = 0.This targets effects exceeding a scientifically substantial magnitude k > 0.
  • Threshold-based identification: With N = 3000, p0 = 0.9, and z_i ∼ N(µ_i,1), the threshold z ≥ 3.5 selects about 63 genes, with P(µ > 2.8|z) = 0.506 at the cutoff.The conditional probability increases as z increases.
  • Threshold-based identification: For the same 63 selected genes, k = 2.0 implies at least a 90% chance at the threshold that µ > 2.0, whereas k = 0 implies about 98% have µ > 0.These posterior probabilities can be averaged across selected cases, analogously to the Benjamini–Hochberg calculation.
  • Decision rules: Researchers can rank genes by P(µ > k|z), choose the number retained, and vary k or the z cutoff according to the desired scientific standard.The proposal assumes a one-tailed setting for large positive effects, while two-tailed probabilities are also available.
  • Model estimation: Estimating g(µ) is generally necessary; exchangeable settings permit estimating f1 and possibly g, whereas nonexchangeable settings make mixing-distribution estimation more difficult.The cited discussion notes that parametric approaches are often easiest in nonexchangeable cases.
  • Applications: The same framework proposes interval-null decisions for hospital profiling, using H1: µ > k to set standards for unacceptable or laudatory departures from average outcomes.Multilevel modeling provides information that independent point-null tests for each hospital would forfeit.

4. NONEXCHANGEABILITY, THE SCHOOL DATA AND THE ONE-GROUP MODEL

The school data are nonexchangeable because school sample sizes vary, and the observed long tails are consistent with that nonexchangeability. The one-group model addresses settings where a sharp null with a substantial point mass is implausible.

  • Nonexchangeability: The 3748 schools are not exchangeable because their sample sizes vary, including different sample sizes across demographic groups within schools.Equal sample sizes could support exchangeability, but the passage says this is rare outside designed experiments.
  • One-group model: The school data show evidence of long tails, with schools having more students tending to be outliers, corresponding to nonexchangeability.The passage links the longer-tailed distribution to variability proportional to sample size.

5. INTERVAL ESTIMATION

The interval-estimation discussion argues that FCR intervals can be too wide because they fail to account for regression toward the mean. It questions a broad conclusion that Bayesian intervals cannot be trusted and emphasizes Bayesian reasoning for constructing intervals with good frequency properties.

  • FCR intervals: In an exchangeable simulation with 1000 random effects, FCR intervals are too wide because their slope is not reduced below 1.0 to track regression toward the mean.A gentler slope near 0.5 can yield shorter intervals with the same coverage rate.
  • Modeling choices: With 1000 observations, exchangeability can support estimating the marginal distribution without assuming Normality, while unequal sample sizes or smaller N may favor Bayesian approaches.The passage notes that nonparametric Bayesian specification is possible but less easy.
  • Interpretive limitation: The discussion calls for more analysis of whether Bayesian reasoning failed in the Section 7 setting and what empirical Bayes adds beyond suggesting frequency-validated Bayesian methods.This marks a scope boundary rather than a resolved conclusion about Bayesian intervals.

6. MODELING AND RTTM

The paper connects multilevel modeling with regression toward the mean and highlights when richer or nonrectangular models are needed. It also describes empirical-Bayes shrinkage of standard deviations as a way to reduce false gene discoveries.

  • Regression toward the mean: Two-level modeling estimates the mean toward which individual random-effect estimates shrink and determines the amount of regression toward the mean.The paper calls the shared information across individuals “ensemble information.”
  • Microarray modeling: Correlations among genes and arrays can improve estimates of f0, f1 and p0 while keeping modeling assumptions to a minimum.
  • Nonexchangeability: Rectangular data representations can obscure nonexchangeability when rows contain different amounts of data, as in the summarized school data.
  • Model assessment: As N increases, richer parametric models can investigate correlations and assess whether exchangeable models are adequate without appealing to asymptotics.
  • Microarray modeling: Shrinking sample standard deviations toward their common mean and using them in t-statistics greatly improves the rate of false gene discoveries.The approach treats the random effects µi and σi as exchangeable and uses empirical-Bayes shrunken estimates.

7. CONCLUSION

The paper presents ideas for analyzing massive datasets while encouraging integration of frequency and Bayesian reasoning with empirical-Bayes modeling. It also emphasizes that exchangeable and nonexchangeable settings require further methodological development.

  • Conclusion: The paper introduces ideas for analyzing massive datasets and encourages a frequency-Bayes unification alongside empirical-Bayes modeling.
  • Conclusion: Its modeling and inference opportunities arise in massive-data settings where exchangeability is assumed.
  • Conclusion: Much remains to understand parametric and nonparametric exchangeable models and to recognize and fit models when nonexchangeability is required.
Loading 0808.0597v1…