Source-linked AI summary
An investigation of the false discovery rate and the misinterpretation of P values
David Colquhoun
TL;DR
The paper examines why conventional P-value interpretation understates the chance that a claimed discovery is false. It uses screening-test analogies and long-run calculations involving power and the prevalence of real effects, finding a 36% false discovery rate in an idealized P = 0.05 example. It concludes that stricter thresholds are needed, while emphasizing that the calculation depends on assumptions about effect prevalence and experimental conditions.
Problem
Biomedical papers commonly misinterpret significance tests, overlooking that the false discovery rate can substantially exceed the nominal P-value threshold.
Method
The paper uses screening-test analogies and long-run conditional-probability calculations based on power and the assumed fraction of tests with real effects.
Results
36% of positive tests are false positives in an idealized example with 10% real effects, 80% power, and P = 0.05.
Takeaways & Limitations
A P value below 0.05 does not by itself ensure a low false discovery rate; the paper discusses P below 0.001 as a stricter criterion.
Takeaways & Limitations
The analysis concerns a single P value under ideal assumptions, and its false discovery-rate calculation requires an assumed fraction of tests with real effects.
Abstract
from arXiv · showhide
The following proposition is justified from several different points of view. If you use P = 0.05 to suggest that you have made a discovery, you will be wrong at least 30 percent of the time. If, as is often the case, experiments are under-powered, you will be wrong most of the time. It is concluded that if you wish to keep your false discovery rate below 5 percent, you need to use a 3-sigma rule, or to insist on P value below 0.001. And never use the word "significant".
Introduction
The paper frames false discovery rate as a neglected contributor to irreproducible biomedical findings and introduces the problem through screening-test examples. Even a test with 80% sensitivity and 95% specificity can produce mostly false-positive results when the condition is uncommon.
- Introduction: The reproducibility crisis is linked to widespread failure to appreciate what governs false discovery rates in biomedical papers.The paper presents this as one contribution to the abundance of irreproducible published findings.
- Introduction: False discovery rate is the probability that a declared significant result reflects chance rather than a real effect.The paper also describes it as the complement of positive predictive value.
- Introduction: 29% is the minimum chance of a false discovery when P = 0.05 defines statistical significance, according to one cited argument.The paper presents this as a reason to question the conventional cutoff.
- Introduction: The screening example is used as an accessible introduction to the analogous hazards of interpreting significance tests.The paper argues that similar reasoning applies to significance tests in biomedical literature.
- Introduction: 86% of positive screening tests are false positives when prevalence is 1%, specificity is 95%, and sensitivity is 80%.Among 575 positive tests, 495 are false positives, leaving an 80/575 probability of actually having MCI of 13.8%.
The significance test problem
The paper distinguishes the classical P value from the long-run probability that a declared effect is real, then calculates false discovery rates using power and the prevalence of real effects. In an idealized example with P = 0.05, the false discovery rate is 36%, and the paper discusses stricter thresholds while noting that these calculations depend on assumptions about effects and testing conditions.
- The significance test problem: A P value of 0.05 means that, if there were no true effect, equally or more extreme data would occur with probability 5%.It does not directly give the probability that a declared effect is real.
- The significance test problem: The paper analyzes a single P value in an ideal case with Gaussian errors, perfect randomization, and one pre-defined outcome, excluding multiple comparisons.The authors state that real experiments can only be less perfect than this setup.
- The significance test problem: 80% power means detecting a real effect in 80% of tests, and power is the significance-test analogue of screening sensitivity.Power depends on sample size and the effect size the experiment aims to detect.
- The significance test problem: The false discovery rate requires an assumed fraction of tests with real effects, which is usually unknown and depends partly on experiment selection.If no tested effects are real, every significant result is a false positive and the false discovery rate is 100%.
- The significance test problem: 36% of positive tests are false positives when 10% of 1000 tests have real effects, power is 80%, and P = 0.05.There are 45 false positives and 80 true positives, for 125 positive tests overall.
A few more complications
Simulations show that interpreting P values requires assumptions about true effects and statistical power: under the null, 5% are false positives, while with a real effect, 78% are detected but selected estimates are inflated. The resulting false discovery rate depends strongly on how common real effects are.
- Interpreting P values: 5% of tests are false positives when the null hypothesis is true and P ≤ 0.05 is used as the threshold.In 100,000 simulations with identical true means, all tests below the threshold are false positives.
- Real effects and power: 78% of tests detect a real difference when 16 observations per group provide power of 0.78.The simulated observed differences are centred around the true difference of 1.
- Real effects and power: The average effect among results with P ≤ 0.05 is 1.14 rather than 1.0 because larger random estimates are more likely to be declared significant.This selection effect makes measured effects too large among positive findings.
- False discovery rate: 36% of positive tests are false positives when 10% of experiments contain real effects.The calculation combines false positives from null cases with true positives from experiments containing genuine effects.
- False discovery rate: The false discovery rate changes with the assumed fraction of experiments having genuine effects, reaching 6% when that fraction is assumed to be 50%.The passage notes that this assumption is uncertain and that underpowered studies can still produce a false discovery rate above 0.05.
Underpowered studies
Underpowered studies detect real effects infrequently, and repeated tests can produce substantially different P values. The paper notes that low power has been common across research fields and remains a practical concern.
- Research practice: Published studies often have power far below 0.8, with values around 0.5 common and 0.2 not rare.The paper attributes low power to small effects, inadequate sample sizes, and ignored statistical warnings.
- Research practice: Median statistical power in neuroscience studies was estimated at about 8% to 31%.The paper describes this range as disastrously low and reports that it has not improved over five decades.
- Power: Power falls from 0.78 with 16 observations per group to 0.46 with 8 and 0.22 with 4.With 4 observations per group, there is only a 22% chance of detecting a real effect when it exists.
- Power: Only 22% of P values are ≤ 0.05 when each group has 4 observations and test power is 0.22.If one test is significant, the next is not significant in 78% of cases.
- Reproducibility: P values are not reproducible under low power, often differing greatly when an experiment is repeated.The simulations show a broad P-value distribution when power is 0.22.
The inflation effect
Selective reporting of P ≤ 0.05 results inflates estimated effects, especially when tests have low power. Simulations also show that false discovery rates rise far above 5%, reaching extreme levels when genuine effects are uncommon.
- An estimated effect of 1.4 at power 0.46 and 1.8 at power 0.22 exceeded the true difference of 1.0.The inflation arises because larger-than-average effects are more likely to cross the significance threshold.
- False discovery rates ranged from 6% to 70% across the examined powers and proportions of experiments with real effects.The problem worsens as the proportion of experiments containing genuine effects decreases.
- 26% of P values between 0.045 and 0.05 were false positives in the most optimistic case with power near 80%.This scenario assumed half of experiments had true effects and half did not.
- When 90% of experiments had no real effect, 76% of just-significant results were false positives, almost independently of power.The large number of false positives can overwhelm true positives when genuine effects are uncommon.
- A P value near 0.05 provides limited evidence of discovery, and these results address only tests close to that threshold rather than lifetime error rates.The simulations therefore do not directly quantify how often a researcher is wrong across all results.
- Berger’s calibration associates P = 0.0027 with a false discovery rate of 0.042, roughly corresponding to a 3-sigma policy.The paper presents P near 0.001 as a stricter threshold associated with a lower false discovery rate.
Is the argument Bayesian?
The paper’s calculations use conditional probabilities and can resemble Bayesian reasoning without requiring subjective probabilities. It connects this framework to practical recommendations for reducing false discoveries and improving statistical reporting.
- The analysis uses conditional probabilities and contains no subjective probabilities, despite its formal similarity to a Bayesian argument.
- A prevalence-based interpretation makes P(real) a long-run probability across many candidate drugs, analogous to selecting active balls from an urn.
- The screening-test analogy shows that false-discovery calculations need not rely on subjective Bayesian probabilities.
- 30% is presented as a minimum false-discovery rate for P = 0.05 under optimistic experimental assumptions.
- The paper recommends reporting P values with effect sizes and confidence intervals, avoiding “significant,” and treating P values near 0.05 as worth another look.
- To reduce false discoveries, the paper recommends P < 0.001 or a three-sigma rule, while warning that underpowered studies inflate false-discovery rates and effect sizes.
APPENDIX
The appendix frames the calculations through standard conditional-probability rules. It explains how joint events and conditional probabilities yield the relevant multiplication relationships.
- The appendix states that the probability of observing events A and B can be expressed in two equivalent ways.
- When A and B are independent, P(A|B) equals P(A), reducing the relationship to the multiplication rule of probability.
The screening example
The screening example demonstrates how prevalence, sensitivity, and specificity determine the probability that a positive result reflects a real condition. A seemingly good test can therefore produce many false discoveries.
- The screening calculation defines illness as event A and a positive test as event B, then seeks the probability of illness given a positive result.
- The false discovery rate is calculated as the probability of a positive test among people without illness, weighted by the non-ill population, divided by all positive-test probability.
- 86.1% is the screening test’s false discovery rate, obtained as 1 − 0.139 from the conditional-probability calculation.
- The paper applies the same screening logic to significance tests by treating a real effect as analogous to prevalence and a significant test as analogous to a positive result.
Bayes’ factor –posterior odds
The paper uses likelihood ratios and Bayes factors to relate observed P values to posterior odds and false-discovery rates. These approaches again show that P = 0.05 can correspond to a false-discovery rate far above 5%.
- Likelihood is the probability of observing data given a hypothesis, and the likelihood ratio compares that probability under H0 and H1.
- The odds on H0 equal the prior odds P(H0)/P(H1) multiplied by the likelihood ratio B.
- 36% is the false discovery rate obtained again from the P = 0.05 example, far above the test’s nominal 0.05 level.
- Berger and colleagues address the unknown prior by deriving a lower bound for the Bayes factor favoring H0 over H1.
- For P = 0.05, the minimum false discovery rate is 0.289 under equal prior probabilities for H1 and H0.
- Table A1 lists P values with corresponding conditional error probabilities calculated from equation A5.