Source-linked AI summary
Terminologies for Reproducible Research
Lorena A. Barba
TL;DR
The paper addresses contradictory terminology surrounding reproducible research, an issue attracting attention across disciplines and stakeholder groups. It inventories terminology across fields using a decision tree and finds distinct camps that reverse the meanings of reproduce and replicate, recommending conversations with organizations and fields to explore alignment.
Problem
Different research communities use reproduce and replicate inconsistently, complicating efforts to address reproducibility concerns.
Method
The paper inventories terminology across publications and classifies it with a decision tree distinguishing undifferentiated usage from two opposing conventions.
Results
The literature falls into a no-distinction group or two camps that assign the minimum same-data, same-methods standard to either reproduce or replicate.
Takeaways & Limitations
Terminology unification should begin with conversations with ACM and FASEB, followed by fields where replication is predominant.
Takeaways & Limitations
The ACM terminology rests on a physical-measurement analogy that is distant from the complex processes of a full scientific workflow.
Abstract
from arXiv · showhide
Reproducible research---by its many names---has come to be regarded as a key concern across disciplines and stakeholder groups. Funding agencies and journals, professional societies and even mass media are paying attention, often focusing on the so-called "crisis" of reproducibility. One big problem keeps coming up among those seeking to tackle the issue: different groups are using terminologies in utter contradiction with each other. Looking at a broad sample of publications in different fields, we can classify their terminology via decision tree: they either, A---make no distinction between the words reproduce and replicate, or B---use them distinctly. If B, then they are commonly divided in two camps. In a spectrum of concerns that starts at a minimum standard of "same data+same methods=same results," to "new data and/or new methods in an independent study=same findings," group 1 calls the minimum standard reproduce, while group 2 calls it replicate. This direct swap of the two terms aggravates an already weighty issue. By attempting to inventory the terminologies across disciplines, I hope that some patterns will emerge to help us resolve the contradictions.
Introduction
Reproducible research has become a cross-disciplinary concern, but efforts to address it are complicated by contradictory uses of key terminology.
- Reproducible research is receiving attention from funding agencies, journals, professional societies, and mass media.
- Different groups use the terms reproduce and replicate in contradictory ways.
- A workshop at the National Science Foundation assessed the variety of reproducibility terminologies.
Pioneers of reproducible research
Early reproducible-research work centered on computational science and argued that readers should be able to rebuild published results from authors’ code, data, and software environments. These efforts focused primarily on computational analyses of unique data rather than studies using newly collected data.
- Claerbout’s group appears to have introduced the phrase “reproducible research” in a 1992 invited paper on computational seismic research.
- The early vision required readers to rebuild published results using the author’s underlying programs and raw data.This implicitly advocated sharing code and data.
- Their reproducibility standard also required a complete software environment and full source code for inspection, modification, and use under varied parameter settings.
- Donoho and colleagues defined reproducible computational research as making all computational details—code and data—conveniently available to others.
- The pioneering work focused on computational analysis of unique seismic data, not on repeating studies with newly collected data.
Conflicting terminologies
Research communities use reproducibility terminology inconsistently: some treat “reproduce” and “replicate” as interchangeable, while others assign them opposing meanings along a spectrum from rerunning analyses to obtaining matching findings with new data.
- The Claerbout/Donoho/Peng convention defines reproducible research as rerunning an analysis with the original data and computer code.
- In the same convention, replication means obtaining the same scientific findings through new data, potentially different methods, and new analyses.
- Drummond (2009) explicitly swapped the prevailing uses of “replicability” and “reproducibility,” without providing much justification.
- The terminology conflict spread across fields, with later publications adopting Drummond’s swapped usage and others reporting directly contradictory definitions.
- The ACM distinguishes repeatability as same-team, same-setup measurement; replicability as different-team, same-setup measurement; and reproducibility as different-team, different-setup measurement.
- These ACM definitions derive from metrology terminology concerning measurement conditions, repeatability, and reproducibility.
- The metrology analogy is limited because physical-measurement terminology does not define “replicability” and may map reproducibility differently across research contexts.
Cataloguing the reproducibility literature
The paper catalogs reproducibility terminology through a decision tree and finds three broad patterns: indistinguishable usage, or two opposing conventions that assign different labels to the same methodological spectrum. The inventory shows that terminology varies across disciplines and remains a practical obstacle to scientific communication.
- Decision-tree classification: Authors either make no distinction between “reproduce” and “replicate,” or distinguish them using one of two opposing conventions.
- Motivation: The catalogue is presented as a response to conflicting terminology that can impede scientific progress and for which defining terms locally remains the available practice.
- Decision-tree classification: The classified spectrum runs from “same data+same methods=same results” to independent studies using new data or methods to obtain the same findings.
- Group B1: Group B1 generally calls rerunning analyses with the same data and code reproducibility, while reserving replication for independent studies or new data.
- Classification limits: Some classifications remain uncertain when papers lack explicit definitions or use terminology unevenly.
- Group B2: Group B2 uses the opposite distinction, calling repeated experiments replicability and new experiments reproducibility.
- Group A: The catalogue includes works that use “reproduce” and “replicate” interchangeably, including major empirical replication efforts.
Expanded terminologies
Disciplines have expanded reproducibility terminology into multiple categories, while economics and political science use replication as an umbrella term. These additions further distinguish computational, empirical, statistical, and inferential dimensions.
- Expanded terminologies: Economics and political science use replication as an umbrella term encompassing pure, statistical, and scientific replication.Pure replication reanalyzes the same data with the same model and parameters; statistical replication may use comparable data; scientific replication uses different samples or populations and possibly different models.
- Expanded terminologies: Neuroscience distinguishes internal, external, cross, and reproducibility categories based on who repeats the analysis, which materials are used, and whether findings are compared.The supplied passage identifies internal and external replicability and indicates further distinctions within the framework.
- Expanded terminologies: Goodman and colleagues propose methods reproducibility, results reproducibility, and inferential reproducibility as a new lexicon for nonstandardized basic terms.Methods reproducibility corresponds to the original Claerbout/Donoho meaning, while results reproducibility corresponds to Peng’s meaning of replicability.
Beyond the academic literature
Agencies, societies, and journals promote transparency and repeated studies, but they use reproducibility and replication inconsistently. Some define the terms distinctly, while others use replication broadly or leave terminology undefined.
- Funding agencies: The NSF SBE report defines reproducibility as duplicating results with the original materials and replicability as producing new evidence through new experimentation.It also calls reproducibility a minimum necessary condition for a finding to be believable and informative.
- Funding agencies: The NSF CISE letter encourages complete protocols, experimental parameters, and collected data but does not define reproducibility, while NIH provides no terminology.NIH instead emphasizes rigor in methods and materials and transparency in publication.
- Professional societies: The AEA uses replication for full-study confirmation and requires clearly documented, readily available data, code, configuration files, and scripts.Its policy applies across all AEA journals and treats replication as repeating a published paper or full study.
- Professional societies: The APA uses replication studies for complete studies, including data collection, intended to confirm another study’s findings; APS promotes replication through reporting and statistical initiatives.APS initiatives include stronger methods reporting and expanded statistical requirements.
- Professional societies: FASEB reverses the common distinction by defining replicability as repeating a specific experiment with the same materials and methodologies, and reproducibility as similar results using comparable materials and methodologies.FASEB limits replicability to a specific experiment rather than an entire study.
- Journals: Journals implement related practices through data archives, replication sections, replication materials, open data, open-source code, and rerunnable analyses.Examples include JAE’s archive and replication section, AJPS replication files, Biostatistics’ reproducibility conditions, Genome Biology’s code requirement, and ReScience’s replication focus.
Conclusions
The Claerbout/Donoho/Peng terminology is widely disseminated, but ACM and FASEB’s opposing adoption makes standardization difficult. The paper proposes conversations and accommodation across disciplines as a feasible path toward unification.
- Conclusions: The Claerbout/Donoho/Peng terminology is broadly disseminated across disciplines, while ACM and FASEB have adopted opposing terminology.The paper describes the justification for ACM’s adoption as tenuous and the source of FASEB’s adoption as unclear.
- Conclusions: Table 2 groups terminologies by discipline, documenting how usage patterns vary across fields.The caption describes the table as grouping terminologies as in Table 1, but by discipline.
- Conclusions: The proposed unification strategy begins with conversations with ACM and FASEB, followed by discussions with political science and economics.The paper suggests absorbing research compendia called “replication files” into a convention centered on reproducible research and open archival practices.