Source-linked AI summary
Impact Factor: outdated artefact or stepping-stone to journal certification?
Jerome K. Vanclay
TL;DR
The paper asks whether the widely used Garfield index and Thomson Reuters Impact Factor remain fit indicators of journal standing. It reviews how the indices are defined, examines literature and database, sampling, and statistical problems, and recommends substantial reform alongside stronger journal certification. The review concludes that the TRIF has serious limitations and should be overhauled or replaced as a proxy for journal quality.
Problem
The TRIF is widely used to assess journal standing despite serious concerns about its reliability as a proxy for journal quality.
Method
The paper reviews the Garfield index and TRIF definitions, synthesizes critical literature, and examines database, sampling, statistical, and text-analytic evidence.
Results
The review finds serious limitations in the TRIF and compelling reasons to overhaul or replace it as a journal-quality indicator.
Takeaways & Limitations
The paper recommends corrected computation and reporting, together with stronger certification of editorial and review procedures.
Takeaways & Limitations
TRIF values depend on database coverage, inclusion policies, citation errors, sampling, skewed citation distributions, and methodological choices.
Abstract
from arXiv · showhide
A review of Garfield's journal impact factor and its specific implementation as the Thomson Reuters Impact Factor reveals several weaknesses in this commonly-used indicator of journal standing. Key limitations include the mismatch between citing and cited documents, the deceptive display of three decimals that belies the real precision, and the absence of confidence intervals. These are minor issues that are easily amended and should be corrected, but more substantive improvements are needed. There are indications that the scientific community seeks and needs better certification of journal procedures to improve the quality of published science. Comprehensive certification of editorial and review procedures could help ensure adequate procedures to detect duplicate and fraudulent submissions.
Abstract
This review examines weaknesses in the Garfield index and its Thomson Reuters implementation, arguing that the TRIF should be substantially improved or replaced as a proxy for journal quality. It also considers text-analytic evidence and the need for stronger journal certification and quality control.
- Scope: The review distinguishes the general Garfield index from the specific Thomson Reuters Journal Impact Factor implementation.The paper adopts Garfield index for the generic concept and TRIF for the Thomson Reuters Journal Impact Factor.
- Database coverage: Different databases produce different citation counts because they scan different source collections and apply different inclusion policies.Scopus, Web of Science, and Google Scholar identify overlapping but nonidentical citations and sources.
- Precision: The TRIF’s three-decimal display suggests more precision than warranted by variation across databases, corrections, methodologies, and journal sizes.The index can vary with database choice, numerator and denominator definitions, and daily error corrections, especially for journals publishing few articles.
- Literature: The literature is predominantly critical of the TRIF and provides compelling reasons to overhaul or replace it as a proxy for journal quality.The review notes continuing publisher reliance alongside substantial academic scepticism and calls for reform.
- Limitations: The TRIF has serious limitations involving database coverage, citation errors, sampling, statistical validity, and unintended consequences.The review groups Garfield-index problems into data errors, system faults, sampling deficiencies, and statistical shortcomings.
Gross Errors
WoS records contain typographic, coding, attribution, and residual test-record errors that complicate reconstruction of the TRIF. These errors make the indicator difficult to reproduce and potentially manipulable.
- Record and citation errors: WoS contained two incorrect records for a Forestry paper, although one record included the correct DOI and citing documents.The errors were attributed to WoS data entry and were quickly amended after being reported.
- Sources and consequences: Citation errors may reflect both author mistakes and database-entry mistakes, with error rates escalating for books and conference proceedings.The paper notes that some errors are faithfully reproduced from citing authors, while others appear to be introduced by Thomson Reuters or its predecessors.
- Record and citation errors: WoS citations can be correct and complete, incomplete, faulty, or linked to ghost documents.These categories represent the four combinations of complete or incomplete and correct or incorrect citation links.
- Reconstructing the TRIF: The 2010 JCR reported 127 citations and a TRIF of 1.460 for Forestry, whereas WoS reconstruction found 143 citations and an impact factor of 1.3 to 1.5 depending on assumptions.The discrepancy reflects problematic citations and uncertainty about which records the TRIF calculation relies upon.
- Record and citation errors: WoS encoded mismatched authors, years, volumes, and journals as citations to Forestry articles.Examples included a paper cited as Šumarstvo being encoded as Forestry and an iForest article being encoded as Forestry.
- Reconstructing the TRIF: WoS errors and subsequent database updates make it impossible to independently reproduce the TRIF or establish whether erroneous citations were omitted.The paper also identifies ghost citations as a possible way to inflate the indicator.
System Faults
System faults in WoS and JCR mishandle journal dates, name changes, and suspensions, producing anomalous records and biased impact-factor calculations.
- Date and journal-history errors: WoS recorded 19 citations to PLOS One articles published in 2005 even though PLOS One began publication in 2006.Some DOI and article-number details could resolve the anomalous publication year.
- Date and journal-history errors: Renaming the Australian Meteorological Magazine produced separate TRIF values of 0.179 and 0.935 instead of a combined value of 0.576.The paper argues that the TRIF should accommodate journal name changes efficiently.
- Date and journal-history errors: The World Journal of Gastroenterology’s suspension caused its 5-year TRIF denominator to exclude 2005 documents while including 2313 citations to them.The paper states that this mismatch biased the TRIF upward.
- Date and journal-history errors: Journal suspensions and reinstatements create inconsistent treatment of documents and citations in JCR calculations.The proposed correction is to include 2005 documents or omit 2005 citations so the calculation compares like with like.
Sampling Deficiencies
The TRIF relies on sampled citation data rather than a census, and its fixed two-year window fits some disciplines poorly because citation accrual varies across fields.
- Sampling basis: WoS samples scientific literature using changing inclusion criteria, so the TRIF is not based on a census.Different providers use different literature samples, producing different interpretations of corresponding impact factors.
- Sampling basis: A journal suspension in WoS could have deflated other gastroenterology journals’ TRIFs by as much as 1%.The example concerns the World Journal of Gastroenterology, for which Scopus recorded over 6000 citations to 2004–05 articles after WoS suspension.
- Temporal window: The TRIF counts citations in one calendar year to work published in the previous two years, but that window is inadequate for some disciplines.The passage states that the two-year interval is appropriate for some fields and inadequate for others.
- Temporal window: Citation accrual differed sharply between Nature and Ecology articles published in 1996, with Ecology reaching its citation mode about a decade later.Figure 12 illustrates the differing citation-accrual trends underlying the concern about a fixed window.
Statistical Shortcomings
The TRIF has statistical and procedural shortcomings that limit its use as a journal-quality indicator, motivating reforms and possibly a certification system focused on publication quality control.
- Statistical shortcomings: Citation distributions are highly skewed, with a few frequently cited articles and many rarely cited articles.Because citation data are not normally distributed, mean-based summaries can poorly represent typical citation patterns.
- Indicator validity: The TRIF lacks transparency, repeatability, and rigour, yet remains widely used as a proxy for journal quality and sometimes article quality.Its continued use is attributed to the lack of a better indicator and to an ill-defined notion of journal value-adding.
- Publication quality control: Certification is a weak link in the scientific communication value chain because referee guidelines rarely address key certification aspects and editors may remain unconcerned about misconduct.Reviewers have also been reported as reluctant to alert editors when they detect suspicious findings.
- Publication quality control: Google Scholar could help detect fraud and plagiarism by showing articles with similar text or images alongside multiple article versions.The proposed additions would be useful when researchers compile reviews and meta-analyses.
- Indicator validity: Good science requires a more proactive editorial role that is not reflected in the TRIF, which the paper describes as unfit for its assumed quality-indicator function.The conclusion places responsibility for future quality science communication with editors if TRIF reform does not occur.
- Reform options: The paper proposes verified citations, like-with-like counting, discipline-sensitive timeframes, appropriate rounding, and confidence intervals as TRIF improvements.It also presents community journal ratings and replacing the TRIF with broader journal certification as alternative paths.
- Reform options: Despite repeated calls for reform, the TRIF remains essentially unchanged apart from a five-year variant and additional metrics.Journals continue to promote increases in the TRIF despite the indicator’s described weaknesses.
- Certification: A certification system could rate editorial efficiency, review rigour, value-adding, and procedures for detecting plagiarism, fraud, and other ethical lapses.The proposed system would assess procedures rather than reward high manuscript rejection rates or penalize journals solely for isolated fraud cases.