Source-linked AI summary
Scientific Utopia: II. Restructuring incentives and practices to promote truth over publishability
Brian A. Nosek, Jeffrey R. Spies, Matt Motyl
TL;DR
Scientific incentives create a conflict between researchers’ rewards and getting scientific results right. The paper proposes transparency-based practice changes to make accountability competitive with publication incentives, arguing that transparency can improve practices even without external scrutiny.
Problem
The paper addresses a conflict of interest in which scientists’ rewards depend on publishing rather than solely on getting results right.
Method
The paper proposes changing scientific practices through transparency and accountability so others can detect errors more readily.
Results
Transparency can improve scientific practices even when no one actually examines them.
Takeaways & Limitations
Making practices transparent increases opportunities for others to detect errors when researchers do not get results right.
Abstract
from arXiv · showhide
An academic scientist's professional success depends on publishing. Publishing norms emphasize novel, positive results. As such, disciplinary incentives encourage design, analysis, and reporting decisions that elicit positive results and ignore negative results. Prior reports demonstrate how these incentives inflate the rate of false effects in published science. When incentives favor novelty over replication, false results persist in the literature unchallenged, reducing efficiency in knowledge accumulation. Previous suggestions to address this problem are unlikely to be effective. For example, a journal of negative results publishes otherwise unpublishable reports. This enshrines the low status of the journal and its content. The persistence of false findings can be meliorated with strategies that make the fundamental but abstract accuracy motive - getting it right - competitive with the more tangible and concrete incentive - getting it published. We develop strategies for improving scientific practices and knowledge accumulation that account for ordinary human motivations and self-serving biases.
A true story of what could have been
A seemingly reliable finding that political extremists perceive shades of gray less accurately than moderates vanished in a high-powered direct replication. The experience exposes how publication incentives can conflict with accuracy and underscores the need to make getting it right competitive with getting published.
- Original finding: Moderates perceived shades of gray more accurately than extremists on the left and right (p = .01).Participants completed a perceptual judgment task matching word shades to a black-to-white gradient.
- Replication: A direct replication with 1,300 participants had .995 power to detect the original effect size, but the effect vanished (p = .59).The replication was conducted while preparing the manuscript.
- Replication: The failed replication was not definitive proof that the original effect was false, but it raised enough doubt for reviewers to recommend against publication.The researchers did not ignore the replication because their lab mates knew it had been conducted, making them accountable.
- Implications: Incentives for publishable results can conflict with incentives for accurate results, encouraging decisions that inflate false results in published literature.The conflict can affect research design, analysis, and reporting.
- Implications: Scientific progress requires making incentives for getting it right competitive with incentives for getting it published.Without such changes, researchers may learn to avoid direct replications after losing exciting results.
How evaluation criteria can increase the false result rate in published science
Publishing strongly influences academic career and institutional evaluation, while scarce journal space and high rejection rates pressure researchers to meet publishing criteria and obtain publishable results. Success therefore depends partly on knowing what is publishable and producing results that satisfy those expectations.
- Publication influences hiring, salary, promotion, tenure, grants, and the evaluation and ranking of departments and universities.
- Publishing expectations extend beyond research-intensive faculty to institutions without graduate programs, graduate students seeking employment, and undergraduates applying to top graduate programs.
- Even competent research, effective analysis, and polished writing do not guarantee publication because peer review remains outside the researcher’s control.
- In social and behavioral sciences, journal rejection rates of 70-90% are common, making acceptance an unusual outcome.
- Scarce journal space pressures authors to meet publishing criteria, so publication success depends partly on social savvy about what is publishable and empirical savvy in obtaining publishable results.
A disconnect between what is good for scientists and what is good for science
Publishing is necessary for research impact but can conflict with accurate knowledge accumulation because publishing and truth are not synonymous. Professional incentives and ordinary motivated reasoning can therefore steer scientists toward publishable conclusions and practices at the expense of accuracy.
- A disconnect between what is good for scientists and what is good for science: Publishing enables research impact, but rewarding publication creates a conflict between scientists’ personal interests and objective knowledge accumulation because published findings need not be true.Scientists have an incentive to publish regardless of whether their findings are true.
- A disconnect between what is good for scientists and what is good for science: Professional success can conflict with practices that improve confidence in findings, producing motivated reasoning that favors desired conclusions over accuracy.The authors distinguish accuracy motives from professional motives and note that strong professional motives can drive reasoning toward career advancement.
- A disconnect between what is good for scientists and what is good for science: Ordinary, accepted research practices can increase false results when directional goals lead scientists to justify career-serving decisions as accuracy-driven.These practices are difficult to set aside precisely because they can be common, accepted, and appropriate in some circumstances.
- A disconnect between what is good for scientists and what is good for science: Motivated reasoning can operate unintentionally: scientists accept confirming evidence uncritically, scrutinize disconfirming evidence heavily, and select flexible analyses that make results more publishable.Complex situations, ambiguous information, and multiple legitimate options make these biases especially influential.
Novelty and positive results are vital for publishability, but not for truth
Scientific publishing norms prioritize novelty and positive results over replication and truth, encouraging practices that increase publishability rather than confidence in claims. These incentives make negative findings and false results less likely to enter or leave the literature.
- Replication and truth: Editors and journals favor new findings over replications because replications are viewed as “not newsworthy” and a “waste of space.”This devalues replication even though replication increases confidence in a claim’s truth value.
- Positive results and truth: Publishing or obtaining a positive result does not establish that an effect is true or indicate its probability of truth.The nominal false-positive rate of alpha = .05 has become a de facto publication criterion.
- Incentive structure: When false results enter the published literature, limited replication incentives and few consequences for error make them difficult to expel.The principal incentive described is publication rather than correction or validation.
- Positive results and truth: More than 90% of psychology publications are positive effects, while negative results are less likely to be pursued, reported, or published.The predominance of positive findings appears longstanding and may be increasing.
- Incentive structure: Novelty and positive-result demands encourage generating new ideas instead of additional evidence, suppressing negative results, and optimizing designs, analyses, and reporting for publishability.These incentives shift effort away from testing previously suggested ideas and toward obtaining publishable positive outcomes.
Practices that can increase the proportion of false results in the published literature
Several defensible research and reporting practices can increase the proportion of false published results, while weak incentives for direct replication allow irreproducible findings to persist. Replication is central to scientific verification, yet often avoided because failures of novelty are deemed unpublishable.
- Practices that can increase the proportion of false results in the published literature: Running many low-powered studies, uncritically accepting successful studies, presenting discoveries as confirmatory tests, and avoiding direct replication can increase false findings.The listed practices exploit chance, asymmetric methodological standards, or the preference for novelty over verification.
- Practices that can increase the proportion of false results in the published literature: Researchers can inflate false results by selectively reporting positive or clean findings, stopping or extending data collection strategically, and reporting only successful variables, analyses, or exclusions.These practices may sometimes be justifiable, but selectively disclosing outcomes or analytic choices increases publishability while reducing validity.
- Practices that can increase the proportion of false results in the published literature: These behaviors are common partly because they can be sensible in exploratory research, such as measuring multiple outcomes when little prior knowledge identifies the likely effect.The appropriate response is disclosure of the design decision and rationale, enabling evaluators to compute confidence accurately and motivating replication.
Strategies that are not sufficient to stop the proliferation of false results
Several proposed remedies are insufficient to curb false results: conceptual replication can confirm but not disconfirm original findings, self-correction may take decades, specialized journals carry low status, and education campaigns have produced little change.
- Conceptual replication: Conceptual replication changes key design operationalizations, so successful replications support an effect while failed replications can be dismissed rather than disconfirming it.It is valuable for abstracting theoretical explanations, especially for unobservable constructs, but cannot replace direct replication.
- The mythology of science as self-correcting: False effects can remain for decades because published findings lack a systemic ethic of confirmation or disconfirmation, and retractions are very rare.Researchers may continue relying on false results even after their falsity becomes known.
- Journals devoted to publishing replications or negative results: Journals devoted to negative results or replications are defined as low-importance outlets, reducing authors’ interest in publishing there.The model is characterized as doomed because it identifies the journal with work that other journals will not publish.
- Education campaigns emphasizing the importance of replication and reporting negative results: Education campaigns have not changed daily practices despite more than three decades of methodological discussion about replication and negative-result reporting.There is little disagreement that the file-drawer effect is harmful, yet the problem persists.
Publishing practices are hard to change because innovative research is more important than
Publishing practices are difficult to change because journals and reviewers generally favor innovative, positive findings over replication and negative results. Blanket replication requirements could impede progress, so better solutions should preserve innovation while rewarding confirmation of existing claims.
- Publishing practices are hard to change because innovative research is more important than: Editors and reviewers usually select innovative findings over replication or negative results when publication space is limited.Journals receive more submissions than they can publish and can choose among many neatly packaged innovative articles.
- Increasing expectations of reviewers to catch motivated reasoning and other signs of false results: Increasing reviewer scrutiny for motivated reasoning and other signs of false results is presented as the most reasonable suggestion.The peer-review process is currently the best available method for identifying potentially false results besides authors’ diligence, although it remains only a partial solution.
- Increasing expectations of reviewers to catch motivated reasoning and other signs of false results: Peer review cannot fully address false results because reviewers are volunteers, miss many errors, and assess summarized reports rather than the underlying research.Reports omit much of the research process, including measures, methods, and analysis strategies, while emphasizing a strong narrative of what readers should learn.
- Raising the barrier for publication: Requiring replications of new findings could raise publication standards, but a blanket policy would lengthen an already difficult process and strain editors and reviewers.The pressure to publish would remain strong, while manuscripts already take years to publish and may be reviewed by multiple journals and teams.
- Raising the barrier for publication: Universal replication requirements could stifle risk-taking because direct replication may be resource-intensive or impossible, discouraging challenging ideas needed for innovation.Innovation requires risk-taking, and innovators can be wrong; the central problem is that false results remain in the literature rather than merely entering it.
- Raising the barrier for publication: The best solutions would encourage innovation and risk-taking while simultaneously rewarding confirmation of existing claims.This approach addresses the persistence of false results without making replication essential for every new phenomenon.
Strategies that will accelerate the accumulation of knowledge
The section proposes improving knowledge accumulation by elevating accuracy and longer-term research impact over publication itself. It recommends practices that promote replication, soundness, openness, accountability, and clearer evaluation of research contributions.
- Promoting and rewarding paradigm-driven research: Paradigm-driven research balances novelty and replication by using established procedures to build knowledge, test validity, and extend findings through mechanisms and boundary conditions.Conceptual replication also tests whether findings generalize theoretically rather than being methodologically idiosyncratic.
- Improving methodological practice: Checklists requiring standard methodological and reporting information could prevent errors, improve disclosure, and address shortcomings in design, reporting, and analysis.The cited review found pervasive shortcomings and concluded that most prediction studies in high-impact journals did not follow current methodological recommendations.
- Metrics to identify what is worth replicating: Replication Value metrics could identify which effects are most worth replicating and guide researchers, reviewers, and others in setting priorities and evaluating research.The section also proposes crowdsourcing replication efforts so individual contributions can accumulate into large-scale investigations.
- Journals with peer review standards focused on the soundness, not importance, of research: Peer review focused on soundness rather than importance would make negative results and replications more publishable without defining the journal by otherwise unpublishable research.Making publication trivial would shift incentives toward evaluating research impact and contribution to cumulative knowledge, while reducing barriers to replications and negative results.
- Accountability and implementation: Shifting incentives toward accountability and reputation management can reward researchers for improving the knowledge-accumulation process, despite messier implementation and execution.The section acknowledges that its proposed practices are idealistic, but maintains that scientific practices can be improved to enhance knowledge building.
- Open data, methods and tools, and workflow: Open data, methods, tools, and workflow would facilitate confirmation, extension, critique, error correction, synthesis, reuse, and credit for contributions beyond individual publications.Registration can reduce discrepancies and address the file-drawer effect by recording what research was conducted even when it is not published.