Source-linked AI summary
Citation Statistics
Robert Adler, John Ewing, Peter Taylor
TL;DR
The report addresses the use and misuse of citation data in assessing research quality, where numerical measures are treated as objective substitutes for complex judgments. It examines citation-based evaluation of journals, papers, and scientists, and concludes that citation statistics provide limited information and should be interpreted cautiously alongside other judgments.
Problem
Citation statistics are widely treated as objective measures of research quality, although citation meanings and their relationship to quality are subjective or insufficiently established.
Method
The report examines how citation data and statistics are used and misused to evaluate journals, papers, and scientists, then analyzes citation meanings and recommends wiser use.
Results
62%: a randomly selected Proceedings paper has at least as many citations as a randomly selected Transactions paper, despite the Proceedings impact factor being half as large.
Takeaways & Limitations
Citation statistics can contribute useful information, but assessments should combine them with other judgments rather than rely on a single coarse measure.
Takeaways & Limitations
Citation data provide only a limited and incomplete view of research quality, and available datasets can be incomplete or inaccurate.
Abstract
from arXiv · showhide
This is a report about the use and misuse of citation data in the assessment of scientific research. The idea that research assessment must be done using ``simple and objective'' methods is increasingly prevalent today. The ``simple and objective'' methods are broadly interpreted as bibliometrics, that is, citation data and the statistics derived from them. There is a belief that citation statistics are inherently more accurate because they substitute simple numbers for complex judgments, and hence overcome the possible subjectivity of peer review. But this belief is unfounded.
EXECUTIVE SUMMARY
Citation statistics can provide useful information, but they offer an incomplete view of research quality and are often misunderstood or misused. The report examines these problems across journals, papers, and scientists, urging assessment with other judgments.
- Citation statistics are not inherently more accurate: improper use or misunderstanding can make them misleading.
- Impact factor is a crude journal-ranking statistic, and substituting journal impact factors for individual paper citations is a pervasive misuse.The impact factor captures only a small amount of the citation distribution and higher-impact journals do not necessarily contain more-cited individual papers.
- The h-index and its variants reduce complex scientist citation records to single numbers while losing crucial information needed for research assessment.
- The validity of citation statistics and their connection to research quality remain insufficiently understood and studied.Existing studies have often focused narrowly on correlations with another quality measure rather than on how to derive useful information from citation data.
- Citation statistics can be valuable and practical, but research quality should not be measured with only one coarse tool.The report recommends tempering citation statistics with other judgments, even though this makes assessment less simple.
INTRODUCTION
The report challenges the growing demand for “simple and objective” citation-based research assessment, arguing that numerical metrics can create an illusory objectivity. It advocates using citation statistics cautiously alongside other judgments because research has multiple goals and citation data provide only partial understanding.
- Governments and institutions seek research-quality assessments to guide decisions about future investments in science.
- The emerging preference for citation-based metrics rests on the belief that simple numbers reduce subjectivity and enable effective comparisons across research.
- That confidence is misplaced because improperly used statistics can mislead, while interpreting citation meaning introduces subjectivity rather than eliminating judgment.
- Citation statistics offer only a partial and sometimes shallow view of research, so they should not be the predominant basis for assessment.
- Because research has multiple goals, its value should be judged by multiple criteria, including citations, peer review, esteem, and grant funding where relevant.
- The report examines citation-data misuse for journals, papers, and people, then urges combining cautious statistical interpretation with other judgments.
RANKING JOURNALS: THE IMPACT FACTOR2
The impact factor is a crude journal-ranking statistic whose interpretation depends on citation windows, database coverage, field-specific citation practices, and journal characteristics. It can support cautious within-field grouping, but not mindless or cross-disciplinary comparisons.
- Definition and coverage: The impact factor is computed from citations to articles published during the preceding two years in Thomson Scientific’s indexed-journal collection.Its database covers more than 9000 journals, while mathematics coverage differs substantially across citation databases.
- Confounding factors: The impact factor is not quite an average citations-per-article measure because some non-substantive items are excluded from its denominator but cited in its numerator.Language and journal type also affect citation counts; review journals often receive substantially more citations.
- Field and time effects: Citation practices differ markedly across disciplines, so impact factors should not be used to compare journals across fields.The appropriate citation window also varies: two years suits some fields, whereas mathematics citations often accumulate later.
- Field and time effects: 5- and 10-year impact factors generally track the 2-year impact factor for the 100 mathematics journals examined.Changing the target-year window nevertheless changes journal rankings, usually modestly except for small journals.
- Appropriate use: The impact factor is crude, not useless, and can initially group journals before other criteria refine rankings and verify that the groups make sense.Annual variation, especially for smaller journals, and small differences may be random; ranking requires caution rather than a single-number judgment.
RANKING PAPERS
Using journal impact factors to rank individual papers, authors, programs, or disciplines is a fundamental misuse of citation statistics. Because paper citation distributions are highly skewed, journal averages can produce frequently incorrect individual comparisons.
- Misuse in assessment: The impact factor is increasingly misused to compare individual papers, people, programs, and disciplines, especially in national research assessments.Reported institutional practices assign rankings, funding, or promotion consequences partly or entirely from journal impact factors.
- Why averages mislead: Highly skewed citation distributions and narrow citation windows create many papers with few or no citations, making journal averages poor substitutes for individual counts.Longer source and target periods increase citation counts and can make journal citation behavior easier to distinguish.
- Journal averages versus papers: Proceedings papers averaged 0.434 citations per article, whereas Transactions papers averaged 0.846, about twice as many.The comparison covered 2381 Proceedings papers and 1165 Transactions papers published during 2000–2004.
- Journal averages versus papers: 62% is the probability that a randomly selected Proceedings paper has at least as many citations as a randomly selected Transactions paper.Thus, the journal-average comparison is wrong more often than right despite the Transactions impact factor being about twice as high.
- Why averages mislead: The impact factor provides surprisingly vague and potentially misleading information about individual papers, so paper assessment should begin with the paper’s own citations.It should not be substituted for individual article citation counts when evaluating authors, programs, or disciplines.
RANKING SCIENTISTS
The report examines h-index and related single-number statistics as attempts to rank scientists, arguing that they discard important information and provide weak grounds for comparison.
- The h-index is the largest n for which a scientist has published n articles, each with at least n citations.
- The h-index was designed to replace publication counts and citation distributions with one number focused on the high-citation tail.
- The m-index divides the h-index by years since a scientist’s first paper, while the g-index captures highly cited papers through cumulative citations.
- Evidence that a high h-index indicates scientific importance is unconvincing because showing high h-indices among Nobel laureates does not establish the reverse probability.
- An h-index of 15 retains only that the top 15 papers have at least 15 citations, discarding the rest of the citation distribution.
- Variants claiming to compare researchers across disciplines or institutions are described as naive attempts to represent complex citation records with one number.
- Assessment should pursue understanding rather than merely making scientists comparable, since national bodies’ use of h-indices for ranking is identified as misuse.
THE MEANING OF CITATIONS
Citation statistics appear objective only if citations have a stable meaning, but the report emphasizes that citation motives and interpretations are complex and only weakly tied to research quality.
- Citation-based assessment depends on interpreting what citations mean, yet proponents of citation statistics do not answer that essential question clearly.
- Citations may signify intellectual debt, influence, value, or impact, but the report presents these meanings as broader than a single objective interpretation.
- Citation scholarship distinguishes reward citations acknowledging intellectual debt from rhetorical citations serving other explanatory or argumentative purposes.
- The report concludes that citation meanings are not simple and citation-based statistics are not as objective as proponents assert.
- Most citations are described as rhetorical, and citation choices can reflect author prestige, relationships, journal availability, or convenience rather than paper quality.
- Even reward citations can reflect motives such as persuasion, negative credit, reader alerts, and social consensus, while some notable results disappear through obliteration.
- Correlation between bibliometric indicators and other evaluations does not by itself establish that citation statistics should replace other assessment methods.
USING STATISTICS WISELY
The report argues that responsible assessment requires scrutiny of data quality, statistical assumptions, uncertainty, confounding factors, and the consequences of rankings rather than reverence for numerical complexity.
- Statistics can become a fetish when people treat them as magical representations that reduce complex realities to simple facts.
- Sophisticated citation algorithms can hide assumptions and make results difficult for most people to evaluate, despite claims that they improve impact-factor rankings.
- The report connects research assessment to broader performance-measurement practice and recommends using statistical expertise alongside other assessment methods.
- Impact-factor and citation analyses depend on selective or inaccurate data, including incomplete journal coverage, author-identification problems, and errors in Google Scholar.
- Faulty data can produce faulty conclusions, making the particular collection of citation data an essential but frequently overlooked assessment choice.
- Rankings often lack prespecified statistical models, give scant attention to uncertainty, and ignore confounding factors such as discipline and article type.
- Research assessment has profound effects on scientists, departments, and disciplines, so the validity and limitations of its tools require careful understanding.
- The report endorses institutional comparison as important when pursued collaboratively, while criticizing simplistic procedures that divert attention and resources from understanding differences.