Source-linked AI summary

The Natural Selection of Bad Science

Paul E. Smaldino, Richard McElreath

arXiv:1605.09511v1physics.soc-phstat.AP

TL;DR

The paper examines how publication-oriented incentives interact with research behavior and methods. Using a dynamical model and evidence concerning statistical power, it argues that these incentives select poorer methods and increasingly high false discovery rates, motivating institutional change.

  • Problem

    Publication pressures can encourage scientists to conduct research in ways that produce publishable results, raising concerns about research design and methodological quality.

  • Method

    The paper defines a dynamical model of research behavior in a population of competing laboratories and considers how research methods propagate.

  • Results

    Selection for publishable results leads to the natural selection of poor methods and increasingly high false discovery rates.

  • Takeaways & Limitations

    Without change, existing incentives lead to degradation, making the future of science a critical question and pointing toward institutional change.

  • Takeaways & Limitations

    Higher statistical power may sometimes increase temporarily but can be quelled by publication bias, and eliminating false positives entirely is likely impossible.

Abstract

from arXiv · show

Poor research design and data analysis encourage false-positive findings. Such poor methods persist despite perennial calls for improvement, suggesting that they result from something more than just misunderstanding. The persistence of poor methods results partly from incentives that favor them, leading to the natural selection of bad science. This dynamic requires no conscious strategizing---no deliberate cheating nor loafing---by scientists, only that publication is a principle factor for career advancement. Some normative methods of analysis have almost certainly been selected to further publication instead of discovery. In order to improve the culture of science, a shift must be made away from correcting misunderstandings and towards rewarding understanding. We support this argument with empirical evidence and computational modeling. We first present a 60-year meta-analysis of statistical power in the behavioral sciences and show that power has not improved despite repeated demonstrations of the necessity of increasing power. To demonstrate the logical consequences of structural incentives, we then present a dynamic model of scientific communities in which competing laboratories investigate novel or previously published hypotheses using culturally transmitted research methods. As in the real world, successful labs produce more "progeny", such that their methods are more often copied and their students are more likely to start labs of their own. Selection for high output leads to poorer methods and increasingly high false discovery rates. We additionally show that replication slows but does not stop the process of methodological deterioration. Improving the quality of research requires change at the institutional level.

1. INTRODUCTION

The paper argues that publication-centered incentives can select and propagate poor research methods without deliberate cheating. It combines evidence on persistent low statistical power with a dynamical model to explain this natural selection of bad science.

  • Motivation: Publication-centered incentives can reward and propagate poor research methods, a process the paper calls the natural selection of bad science.The process requires no conscious strategizing or cheating; methods that lead to publication are positively selected.
  • Cultural transmission: Research methods satisfy the conditions for cultural selection because practices vary, affect publication and careers, and are transmitted through mentors, peers, students, and successful role models.Students can carry methods into new laboratories, while prestige-biased adoption spreads practices indirectly.
  • Evidence and approach: The paper supports its argument with empirical and analytical evidence, including a literature review focused on statistical power.The authors examine repeated calls for improved methodology and use statistical power as a concrete case.
  • Evidence and approach: Despite more than 50 years of warnings about low statistical power, the review finds no detectable increase.The authors present this persistence as evidence that poor methods may be sustained by institutional incentives rather than misunderstanding alone.
  • Formal model: A dynamical model of competing, high-integrity research agents shows how publication-linked hiring and retention can explain the persistence of poor practice.The model allows methodology to vary and evolve through consequences for successful publication, even when researchers do not directly respond to incentives for poor methods.

2. INSTITUTIONAL INCENTIVES FOR SCIENTIFIC RESEARCHERS

Academic competition makes publication quantity, novelty, and positive results important career signals. Because positive findings are easier to publish and publication metrics guide evaluation, these incentives can favor methods that generate more publishable results regardless of truth.

  • Institutional incentives: Competition for jobs, grants, promotions, prestige, and graduate-student placement creates incentives for scientists to stand out through output and impact.The paper notes that tenure-track positions and other academic rewards remain highly competitive.
  • Novelty and output: Pressure for novelty is reflected in a 2500% or more increase in the words “innovative,” “groundbreaking,” and “novel” in PubMed abstracts between 1974 and 2014.The authors interpret this language change as a response to increasing pressure for novelty and distinction.
  • Institutional incentives: Publication quantity is often treated as a benchmark of researcher quality, although the exact evaluation standard varies across disciplines.The argument applies wherever the number of published papers, however scaled, is used to assess success.
  • Metric incentives: When evaluative metrics become targets, researchers can exploit them, including through self-citation and fabricated citation networks that increase h-indices.The paper uses these examples to illustrate how incentives can propagate strategies beyond deliberate individual cheating.
  • Selection pressures: Rewarding publication quantity can create selection pressures for cultural evolution of methods that produce more publishable results without requiring direct strategizing.The paper links these pressures to the ecology and transmission of practices within scientific communities.
  • Publication bias: Positive results are more publishable and prestigious than negative results, while failed hypotheses may go unsubmitted or failed replications unpublished.These publication patterns favor researchers who obtain more positive results, whatever their truth value.

3. CASE STUDY: STATISTICAL POWER HAS NOT IMPROVED

Statistical power is a case study of how poor methods can persist despite longstanding recognition of their costs. A review of six decades of evidence finds very low power and no detectable improvement, while publication incentives may favor low-powered studies.

  • Why power matters: Low statistical power increases false negatives and can also increase false discovery rates and inflate reported effect sizes.Its reduced ability to mute stochastic noise contributes to these problems.
  • Incentives and power: Under publication incentives favoring positive findings, low-power studies can be cheap and repeatedly searched for significant results.The paper presents this as a potential conflict between individual publication success and the population’s benefit from high power.
  • Evidence over time: Earlier meta-analyses found low statistical power and no discernible improvement after Cohen’s 1962 warning, while later evidence suggested power remained low.The authors use this history to motivate an expanded review.
  • Results: 0.24 was the mean statistical power, and R2 = 0.00097 indicated no sign of increase over six decades.At this power, tests fail to detect small effects when present three times out of four.
  • Limitations: The evidence is incomplete and may overestimate average power because it draws only on published results that passed peer review.The authors therefore describe the empirical evidence as persuasive but not conclusive.

4. AN EVOLUTIONARY MODEL OF SCIENCE

The paper extends an evolutionary model to heterogeneous laboratories whose research methods are culturally transmitted through differential publication success and lab reproduction. Its assumptions link power, false positives, effort, productivity, replication, and methodological inheritance.

  • Model foundation: The model extends a previous framework to a finite, heterogeneous population of N laboratories investigating novel and previously tested hypotheses.Labs investigate hypotheses experimentally and attempt peer-reviewed publication; the model has Science and Evolution stages.
  • Lab characteristics: Each lab’s power reflects the entire chain of inference and is its probability of correctly detecting a true hypothesis.Power is treated as a characteristic of methodology rather than only of a statistical procedure.
  • Trade-offs: Increasing power also increases false positives unless effort is applied, while greater effort permits better methods but reduces productivity because rigorous research takes longer.The model treats effort as independent of power and uses it to represent rigor in experimental and statistical methods.
  • Evolutionary transmission: Positive results are easier to publish, and publication rewards make more productive laboratories more likely to propagate their methods through new labs and progeny.New labs resemble but are not identical to their parent laboratories, allowing methods to be culturally transmitted with variation.
  • Evolutionary transmission: More successful labs produce more progeny, which inherit their methods as existing labs cease producing research and are replaced within the population.This evolutionary stage makes publication-linked success determine which methodological characteristics become more common.

5. SIMULATION RESULTS

The simulations show that publication-linked selection can drive power, false positives, and effort toward pathological extremes, while replication slows but does not prevent methodological deterioration.

  • Model setup: The model is simplified but isolates how scientific-community dynamics can select research methods over time.The authors introduce the dynamics in stages so readers can examine the isolated forces.
  • 5.1. The natural selection of bad science: Higher power increases positive and publishable results, causing both power and the false positive rate to approach unity.In the simulations, all results eventually become positive and therefore publishable.
  • 5.1. The natural selection of bad science: Reducing effort increases publication rates at the cost of more false discoveries, selecting for continued degradation when negative findings are difficult to publish.The result links output incentives to lower effort and poorer scientific practices.
  • 5.2. The ineffectuality of replication: Replication evolved slowly to around 0.08 because it was guaranteed publication but worth only half as much as a novel result.Replication was therefore only weakly selected for under the model’s incentives.
  • 5.2. The ineffectuality of replication: Low effort increased false positives enough that pursuing novel hypotheses became more lucrative than increasing replication.The model therefore continued selecting low effort despite replication’s availability.
  • 5.2. The ineffectuality of replication: Even replication rates as high as 50% only slowed the decline of effort and did not prevent effort from eventually bottoming out.This result held when replication could not mutate and began at very high levels.

6. DISCUSSION

The discussion argues that publication-centered incentives can culturally select methods that maximize publishable results rather than research quality, increasing false discoveries. It concludes that improving science requires institutional changes that reward quality and reproducibility, although implementing such changes is difficult.

  • Incentives drive cultural evolution, so publication quantity can select methods that produce more publishable results rather than better science.The authors frame this process as occurring without deliberate cheating or loafing by scientists.
  • Poor methods are naturally selected under publication-focused incentives, leading to increasingly high false discovery rates.
  • Shallow work that generates more publications may be favored over difficult research requiring years to produce coherent results, disadvantaging researchers pursuing complex questions.
  • Collaboration can provide higher-quality research as a public good without necessarily improving researchers’ fitness under publication-based cultural selection.
  • Institutional change is needed because bottom-up solutions are unlikely to suffice, yet coordination makes such change costly and difficult for early adopters.
  • Replication and stricter publication standards may limit bad science, but replication cannot eliminate false positives and is difficult or impossible in some fields.
Loading 1605.09511v1…