Source-linked AI summary
Abandon Statistical Significance
Blakeley B. McShane, David Gal, Andrew Gelman, Christian Robert, Jennifer L. Tackett
TL;DR
The paper examines how NHST and threshold-based p-value rules create problems for replication and scientific reasoning. It proposes abandoning statistical significance as the default while treating p-values as continuous evidence alongside domain-relevant factors. The authors present this as a principled change, while acknowledging that their recommendations alone will not resolve the replication crisis.
Problem
NHST’s threshold-based treatment of statistical significance creates problems for replication and scientific reasoning in biomedical and social sciences.
Method
The authors propose replacing NHST’s default threshold-screening role with continuous consideration of p-values alongside currently subordinate factors.
Results
The paper concludes that abandoning statistical significance and p-value thresholds is preferable to treating thresholds as default decision rules.
Takeaways & Limitations
P-values should remain available but should not be thresholded or prioritized over broader evidence in publication and statistical decision making.
Takeaways & Limitations
The authors acknowledge that their recommendations will not themselves resolve the replication crisis.
Abstract
from arXiv · showhide
We discuss problems the null hypothesis significance testing (NHST) paradigm poses for replication and more broadly in the biomedical and social sciences as well as how these problems remain unresolved by proposals involving modified p-value thresholds, confidence intervals, and Bayes factors. We then discuss our own proposal, which is to abandon statistical significance. We recommend dropping the NHST paradigm--and the p-value thresholds intrinsic to it--as the default statistical paradigm for research, publication, and discovery in the biomedical and social sciences. Specifically, we propose that the p-value be demoted from its threshold screening role and instead, treated continuously, be considered along with currently subordinate factors (e.g., related prior evidence, plausibility of mechanism, study design and data quality, real world costs and benefits, novelty of finding, and other factors that vary by research domain) as just one among many pieces of evidence. We have no desire to "ban" p-values or other purely statistical measures. Rather, we believe that such measures should not be thresholded and that, thresholded or not, they should not take priority over the currently subordinate factors. We also argue that it seldom makes sense to calibrate evidence as a function of p-values or other purely statistical measures. We offer recommendations for how our proposal can be implemented in the scientific publication process as well as in statistical decision making more broadly.
1 The Status Quo and Two Alternatives
The authors argue that NHST’s 0.05 threshold gives p-values priority over broader evidence, contributing to replication problems. They propose abandoning statistical significance as the default paradigm while retaining p-values as continuous, non-exclusive evidence.
- Status quo: Published findings are commonly required to reach p < 0.05 before other evidence receives serious consideration.Subordinate factors include prior evidence, mechanism plausibility, study design, data quality, costs and benefits, and novelty.
- Status quo: Statistical significance can arise from pure noise, making low replication rates predictable under existing scientific practices.The authors connect this problem to prominent examples and theoretical work.
- Authors’ proposal: They propose dropping NHST and intrinsic p-value thresholds as the default paradigm for research, publication, and discovery.The proposal applies to biomedical and social sciences.
- Authors’ proposal: P-values should be treated continuously and considered alongside currently subordinate factors as one among many pieces of evidence.The authors do not seek to ban p-values or other purely statistical measures.
- Implementation: The paper recommends implementing this evidence-centered approach in scientific publication and broader statistical decision making.The authors also argue that evidence should seldom be calibrated as a function of p-values alone.
Testing
The authors contend that NHST is poorly suited to biomedical and social sciences because small, variable effects and systematic errors make its sharp null implausible. Thresholded evidence further encourages dichotomous reasoning and misleading interpretations of p-values.
- 2.1 Preface: Biomedical and social-science studies commonly involve small, variable effects, noisy measurements, and systematic errors.Examples of systematic error include measurement problems, biased samples, missingness, non-response, and confounding.
- 2.2 Implausible Null Hypothesis: The conventional null of zero effect and zero systematic error is therefore generally implausible and uninteresting.The authors note that effects would vary across people and contexts even when a phenomenon’s overall effect were zero.
- 2.2 Implausible Null Hypothesis: Lexicographic publication rules can favor noisy, smaller studies that attain significance over fewer, larger, better studies.Significant noisy estimates may be upwardly biased, have the wrong sign, and be encouraged by multiple-comparison practices.
- 2.3 Categorization of Evidence: NHST dichotomizes evidence into significant and nonsignificant categories despite evidence varying continuously.The authors describe the conventional threshold as arbitrary and argue that dichotomization encourages dichotomous thinking.
- 2.3 Categorization of Evidence: The authors argue that p-values are poor evidence measures because they depend on an implausible null and assumptions that often fail.A small p-value may signal a problem with at least one assumption without identifying which one; a large p-value may reflect an insensitive test or offsetting errors.
- 2.4 Erroneous Scientific Reasoning: Researchers often focus on whether p-values fall below 0.05, interpret evidence dichotomously, and confuse statistical with practical significance.The paper also reports widespread misunderstanding of basic p-value statements among students and professors.
3 Problems Specific to the Benjamin et al. (2018) Pro-
The authors argue that lowering the significance threshold to 0.005 does not resolve key replication and inference problems. They identify order dependence, model misspecification, mismatched decision costs, and uncertain consequences as reasons for rejecting threshold-based reform.
- Replication: Under the 0.005 proposal, replication outcomes can depend on which study was conducted first.A study with p < 0.005 and another with p between 0.005 and 0.05 can receive different replication classifications depending on study order.
- Validity: Preregistration does not ensure valid p-values when the model generating them is importantly misspecified.The authors present model misspecification as an additional problem beyond uncorrected multiple comparisons.
- Consequences: The authors do not know whether implementing the 0.005 threshold would improve or degrade science.They identify possible benefits alongside greater overconfidence, exaggerated effect sizes, and discounting of important findings that miss the threshold.
- Conclusion: They conclude that thresholds based on p-values or other purely statistical measures are a bad approach.This conclusion motivates their broader proposal to abandon statistical significance rather than select a different cutoff.
4 Abandoning Statistical Significance
The authors argue that threshold-based statistical significance turns uncertainty into binary claims and should be replaced by continuous, holistic evidence assessment. They recommend demoting p-values while incorporating domain-relevant factors into research, publication, and decision making.
- Problems with statistical significance: Modified p-value thresholds, confidence intervals, and Bayes factors do not resolve the broader problems of threshold-based evidence classification.These approaches can retain thresholding and focus on purely statistical measures rather than holistic evidence.
- Problems with statistical significance: Thresholding p-values and related statistical measures encourages binary declarations of “an effect” or “no effect” from uncertain data.The authors characterize this as uncertainty laundering and statistical alchemy rather than a reliable path to certainty.
- The proposed alternative: The proposal is to treat p-values continuously as one piece of evidence alongside prior evidence, mechanism plausibility, study design, data quality, costs, benefits, and novelty.The authors explicitly reject banning p-values; they reject giving them a threshold screening role or priority over subordinate factors.
- Scope and decision making: The recommendations should vary across domains, and the authors reject a universal template because formulaic application could reproduce the problems of rote NHST.They nevertheless offer broad principles and note that non-statistical thresholds may sometimes be useful in regulatory, policy, or business decisions.
- Implementation in publication: Authors should use these factors to motivate research and analysis, report all relevant data and results, and avoid focusing on isolated comparisons that cross statistical thresholds.Editors and reviewers should evaluate both statistical measures and the broader factors supporting a paper.
5 Discussion
The paper reiterates its proposal to abandon statistical significance, retain p-values as continuous evidence rather than thresholds, and implement this approach in publication and statistical decision making.
- The authors propose abandoning statistical significance and provide recommendations for implementing this proposal in scientific publication and statistical decision making.
- The authors emphasize that p-values are not to be banned, but integrated with broader evidence rather than used as a dominant decision rule.
- P-values and other purely statistical measures should not be thresholded or given priority over currently subordinate factors.
- The proposal extends beyond continuous p-value interpretation by arguing that evidence should seldom be calibrated as a function of purely statistical measures.
- The proposal builds on a longstanding literature opposing threshold-based statistical significance, while claiming additional contributions in prioritizing subordinate factors and offering implementation guidance.
- The recommendations are not expected to resolve the replication crisis themselves, but may redirect researchers toward theory, mechanism, measurement, and acceptance of uncertainty and variation.
Tests
The paper argues that proposed statistical fixes retain fundamental problems because threshold-based testing does not represent scientific learning or context-dependent costs and benefits. In particular, the proposed 0.005 threshold and UMPBT framework introduce additional arbitrary or restrictive assumptions rather than resolving these issues.
- UMPBTs rely on restrictive assumptions, including minimax priors that do not represent effect-size distributions and difficulties with nuisance parameters and complex null hypotheses.
- NHST’s binary decisions and 0-1 loss function do not generally map onto scientific learning or the costs and benefits of outcomes.
- The relevant tradeoff should depend on the costs, benefits, and probabilities of all outcomes, which vary across studies and domains.
- Other criticisms are that UMPBTs are not readily applicable to multivariate or realistic settings and can lose their Bayesian justification for the proposed threshold.
- The 0.005 rule remains an arbitrary threshold because it encodes an implicit false-positive/false-negative tradeoff that is not absolute or context-independent.
B Case Study
The sodium-and-blood-pressure case study illustrates how evidence should be assessed without thresholded statistical significance. Interpretation depends on prior evidence, mechanism, study design, data quality, effect magnitude, uncertainty, costs and benefits, and context, producing a more nuanced account than binary declarations.
- Case Study: A p-value of 0.001 supports the sodium hypothesis, but conclusions about association or causation depend on study design and data quality.
- Case Study: Study findings may not generalize across populations because genetic or dietary differences can change whether an association holds.
- Case Study: Causal support increases when prior studies, physiological and animal evidence, or randomized sodium levels consistently support an effect on blood pressure.
- Case Study: Clinical significance depends on effect magnitudes, downstream outcomes, uncertainty, and intervention costs rather than on a p-value.
- Case Study: Statistical measures should be treated continuously and considered as one among many pieces of evidence without priority over other factors.
- Case Study: A multilevel model can yield estimates varying across health, dietary, demographic, and geographic factors instead of one binary effect/no-effect declaration.
- Case Study: Accepting uncertainty and variation in effects produces a richer and more nuanced account of the sodium–blood-pressure association.