Source-linked AI summary
Why we (usually) don't have to worry about multiple comparisons
Andrew Gelman, Jennifer Hill, Masanao Yajima
TL;DR
The paper addresses the problem of statistical inference across many tests, where classical approaches emphasize Type I error and multiple-comparisons corrections. It proposes Bayesian multilevel models that represent the questions jointly and use partial pooling. The authors report that this approach yields more reliable estimates and comparisons while avoiding the power loss associated with many traditional corrections.
Problem
Testing many hypotheses can increase erroneous rejections, motivating classical multiple-comparisons corrections.
Method
The paper proposes Bayesian multilevel models that represent multiple research questions jointly and use partial pooling to incorporate multiplicity into the model.
Results
Multilevel modeling yields more reliable estimates and comparisons, with the authors reporting better results than the simplest multiple-comparisons approach.
Takeaways & Limitations
For social science and program evaluation, the paper recommends multilevel modeling rather than classical multiple-comparisons corrections.
Takeaways & Limitations
Multilevel modeling can be more challenging for complicated structures, and the authors state that more research is needed.
Abstract
from arXiv · showhide
Applied researchers often find themselves making statistical inferences in settings that would seem to require multiple comparisons adjustments. We challenge the Type I error paradigm that underlies these corrections. Moreover we posit that the problem of multiple comparisons can disappear entirely when viewed from a hierarchical Bayesian perspective. We propose building multilevel models in the settings where multiple comparisons arise. Multilevel models perform partial pooling (shifting estimates toward each other), whereas classical procedures typically keep the centers of intervals stationary, adjusting for multiple comparisons by making the intervals wider (or, equivalently, adjusting the $p$-values corresponding to intervals of fixed width). Thus, multilevel models address the multiple comparisons problem and also yield more efficient estimates, especially in settings with low group-level variation, which is where multiple comparisons are a particular concern.
1 Introduction
Multiple comparisons arise when researchers test many questions, and classical procedures address the resulting error concerns through corrections. The paper instead argues for coherent Bayesian multilevel models that represent these questions jointly, use partial pooling, and can produce more reliable comparisons without the same loss of power.
- Applications and evaluation: The paper illustrates the argument with examples involving multiple policy interventions, social indicators, subgroups, and program-evaluation comparisons.It also aims to demonstrate the procedure’s effectiveness in realistic settings through small examples.
- The multiple comparisons problem: Testing many hypotheses raises the probability of erroneous statistical significance, creating a serious concern for classical inference.The concern applies both when no effects exist and when some true effects coexist with additional non-real significant findings.
- The paper’s perspective: The paper argues that the central problem is insufficient modeling of relationships among parameters rather than multiple testing itself.It questions the Type I error paradigm because strictly true null hypotheses are rarely considered plausible.
- The proposed approach: Partial pooling shifts point estimates and intervals toward each other, whereas classical corrections generally keep estimates stationary and widen intervals or adjust corresponding p-values.The contrast is between changing estimates through shrinkage and changing interval width while leaving centers fixed.
- Implications: Multilevel comparisons are more likely to include zero, making them more conservative while remaining more likely to be valid.The paper presents this as an adjustment that does not sap power to detect true differences as many traditional methods do.
- The proposed approach: Bayesian multilevel models represent relevant research questions as parameters in one coherent model and incorporate multiplicity from the outset.This shifts the burden from post hoc correction to modeling the relationships among parameters.
2 Multiple comparisons problem from a classical perspective
Multiple site-specific tests create a substantial risk of false discoveries, while classical corrections reduce that risk by widening intervals or tightening p-value thresholds. Bonferroni therefore trades fewer false rejections for lower power, whereas FDR procedures offer a less conservative alternative.
- The multiple-comparisons problem: The IHDP analysis examines site-specific treatment effects because participating children and program implementation differed across sites.The experiment randomized within site and birth-weight group, with eight sites treated as blocks for exposition.
- The multiple-comparisons problem: Eight site-specific tests at the 0.05 level create a 34% chance that at least one test rejects erroneously.For two independent tests, the corresponding probability is 1−0.95×0.95=0.098≈0.10.
- Bonferroni correction: Bonferroni divides each original p-value by the number of tests, yielding a 0.05/8=0.0062 threshold for eight tests.The same adjustment can be represented by widening confidence intervals while leaving point estimates unchanged.
- Bonferroni correction: Uncorrected intervals reject no-effect for 7 of 8 sites, compared with 5 sites for multiple-comparisons-adjusted intervals.The adjusted intervals preserve the classical point estimates but widen their uncertainty intervals.
- Bonferroni correction: Bonferroni reduces false rejections at the expense of Type 2 error, potentially reducing power to detect important effects.The correction changes the rejection threshold or widens intervals, increasing cases where a true null rejection is missed.
- Alternatives to familywise-error control: False-discovery-rate control is less conservative than familywise-error control and is more powerful for detecting real effects.The authors note that FDR methods may be less useful in social-science settings with fewer tests and less clearly zero effects.
3 A different perspective on multiple comparisons
The paper challenges classical multiple-comparisons corrections and argues that Bayesian multilevel modeling offers a different perspective by modeling related effects jointly. Partial pooling draws estimates toward one another, reducing misleading sign and magnitude inferences while reflecting group-level information.
- A different perspective on multiple comparisons: Classical procedures focus on Type I errors, but the paper argues they insufficiently model the ensemble of parameters underlying related tests.The authors present a different Bayesian perspective rather than proposing an optimal method for all circumstances.
- A different perspective on multiple comparisons: Type S errors occur when researchers infer the wrong sign or ordering of effects, while Type M errors concern substantially misjudging effect magnitude.Examples include claiming a beneficial effect is harmful, reversing site comparisons, or calling a near-zero effect large.
- A different perspective on multiple comparisons: With standard deviation 3 rather than 1, an estimator is more likely to produce a large estimate when the true effect is zero.Higher uncertainty therefore increases the probability of Type M errors, including when examining subgroup rather than main effects.
- Multilevel modeling in a Bayesian framework: Bayesian multilevel models use partial pooling, compromising between complete pooling and separate site estimates by allowing treatment effects to vary across sites.The model can also pool effects toward a fitted group-level regression surface when group-level predictors are included.
- Multilevel modeling in a Bayesian framework: Partial pooling shifts point estimates toward one another instead of merely widening uncertainty intervals as classical multiple-comparisons procedures do.This movement reflects information about the effects and the main effect across sites rather than inflating uncertainty estimates.
- Multilevel modeling in a Bayesian framework: The amount of shrinkage changes when Site 3 is subsampled or bootstrapped because its estimate’s uncertainty relative to the grand mean changes.Overall uncertainty increases when the model has less certainty about treatment-effect heterogeneity across sites.
- Multilevel modeling in a Bayesian framework: Partial pooling reduces comparison z-scores because posterior means are pulled together faster than posterior standard deviations decrease.The z-score reduction becomes stronger as group-level variance approaches zero, and statistically significant Bayesian comparisons become less likely.
4 Examples
The examples argue that classical multiple-comparisons corrections can be inappropriate when true differences are not exactly zero, while multilevel models adapt uncertainty to the setting. In simulations and applications, hierarchical shrinkage produces more reliable and often more informative comparisons.
- Comparing average test scores across states: Classical multiple-comparisons procedures can be inappropriate when the null hypothesis of exactly zero differences is false.The paper argues that such procedures continue widening intervals as comparisons are added, ignoring information in the data.
- Comparing average test scores across states: Multilevel modeling can directly summarize information relevant to both interval coverage and Type S error concerns.The paper recommends modeling raw state averages, predictors, and scores from other years when available.
- Comparing average test scores across states: Multilevel models provide more informative state comparisons, with more claims made confidently and fewer ambiguous comparisons.In the NAEP example, the classical procedure overcorrects when true differences between states are large.
- SAT coaching in eight schools: When group-level variation is low, hierarchical Bayesian estimates pool strongly toward a common mean, effectively performing a multiple-comparisons correction.In the eight-schools example, the estimated group-level variance is zero and none of the Bayesian comparisons is close to statistically significant.
- Fishing for significance: For the beautiful-parents example, the study’s low sample size and noisy, poorly defined predictor limit conclusions beyond a possible overall trend.The paper states that the sample is too small given known risks of Type S and Type M errors.
5 Multiple outcomes and other challenges
Multiple outcomes can be modeled hierarchically by allowing treatment effects to vary across sites and tests while incorporating outcome structure. The approach becomes more complex for richly structured comparisons, where the model must match the problem structure.
- Multiple outcomes: A simple multilevel model can accommodate multiple outcomes measured across related domains or time points.Modeling trends over time may require a more complicated strategy, but is otherwise described as a reasonable choice.
- Multiple outcomes: The IHDP multiple-outcomes model allows treatment effects to vary by site and test across eight cognitive outcomes.The outcomes were measured at ages 3, 5, and 7, with site- and test-specific effects.
- Multiple outcomes: The model additionally allows test effects to vary systematically by test age and verbal versus performance skills, supporting its exchangeability assumption.Test scores were standardized within the sample before fitting the model.
- Multiple outcomes: Across the IHDP outcomes, treatment effects were larger on average for tests taken immediately after the intervention, with similar patterns across sites.The authors note that site-by-outcome interactions could relax the assumption underlying these similar patterns.
- Multiple outcomes: With 64 comparisons, Bonferroni intervals become more uncertain, whereas multilevel estimates shrink toward the grand mean and have vastly greater overall precision.The comparison includes eight outcomes across the eight sites.
- Further complications: More complex comparison structures require further work so the model matches their structure rather than treating all combinations as exchangeable groups.The paper gives a 2 × 3 × 4 × 5 example that should not be modeled as 120 exchangeable groups.
6 Conclusion
The authors argue that multilevel modeling is preferable to classical multiple-comparisons adjustments because partial pooling yields more reliable estimates without simply widening intervals. They recommend extending multilevel models while recognizing challenges for complicated structures and the need to map methods to applied settings.
- Multiple comparisons can create problems, but the authors prefer multilevel modeling over methods that alter p-values or widen confidence intervals.They frame the issue in terms of Type S or Type M errors rather than assuming effects are exactly zero.
- Partial pooling shifts point estimates and corresponding intervals closer together where necessary, especially when much variation is attributable to noise.This approach is presented as yielding more reliable estimates.
- The authors acknowledge that multilevel modeling can be more challenging for complicated structures and identify this as an area needing further research.They also note that fitting functions are available in many statistical software packages.
- They argue that research effort should prioritize expanding the multilevel-model framework rather than classical adjustments based on a perspective they reject.Their recommendation is explicitly tied to the applied setting.
- Multilevel models are argued to improve on simple multiple-comparisons corrections without imposing more burden than sophisticated classical corrections.The authors note that multilevel models may be especially appropriate for clustered studies because they reflect within-group error correlation.