Source-linked AI summary
A review of the literature on citation impact indicators
Ludo Waltman
TL;DR
Citation impact indicators are increasingly important in research evaluation, but their extensive literature requires synthesis. This paper conducts an in-depth review spanning databases and selected issues in indicator construction, concluding with recommendations for future research. The review also highlights limits of journal-level indicators and unresolved consequences of construction choices.
Problem
Citation impact indicators play an important role in research evaluation, creating a need to examine the literature on their construction and use.
Method
The paper provides an in-depth review of citation-impact literature, including databases and selected topics such as counting methods and journal indicators.
Results
97% of publications from 2005 covered by Web of Science are also covered by Scopus.
Takeaways & Limitations
Journal-level citation impact indicators should not substitute for publication-level citation statistics.
Abstract
from arXiv · showhide
Citation impact indicators nowadays play an important role in research evaluation, and consequently these indicators have received a lot of attention in the bibliometric and scientometric literature. This paper provides an in-depth review of the literature on citation impact indicators. First, an overview is given of the literature on bibliographic databases that can be used to calculate citation impact indicators (Web of Science, Scopus, and Google Scholar). Next, selected topics in the literature on citation impact indicators are reviewed in detail. The first topic is the selection of publications and citations to be included in the calculation of citation impact indicators. The second topic is the normalization of citation impact indicators, in particular normalization for field differences. Counting methods for dealing with co-authored publications are the third topic, and citation impact indicators for journals are the last topic. The paper concludes by offering some recommendations for future research.
1. Introduction
Citation impact indicators have become increasingly important in research evaluation, motivating a rapidly expanding literature. This paper reviews that literature in depth, covering databases, indicator construction choices, selected topics, and future research recommendations.
- Citation impact indicators have gained importance in research evaluation, alongside a rapidly growing research literature.
- The paper presents an in-depth, large-scale review intended for researchers and practitioners working with citation impact indicators.
- Selected topics include publication and citation selection, normalization, counting methods, and citation impact indicators for journals.
- It reviews literature on Web of Science, Scopus, and Google Scholar as databases used to calculate citation impact indicators.
- The review excludes several areas, including detailed coverage of the h-index, interpretation, practical research-evaluation applications, peer-review correlations, and historical development.
- The paper concludes with recommendations for future research.
2. Methodology
The review used a semi-systematic, citation-network-based process to identify and assess literature on selected citation-impact topics. It searched core and related journals, expanded candidate publications through citation links, and manually judged relevance, while remaining selective rather than exhaustive.
- Literature identification: The literature was collected through a semi-systematic methodology combining prior knowledge, preliminary searches, and citation-network exploration.CitNetExplorer was used to identify additional publications connected to an initial relevant set.
- Literature identification: The network was expanded by identifying cited, citing, or closely connected publications and manually checking their titles and abstracts for topic relevance.Several iterations were performed until the relevant literature on a topic seemed to have been found.
- Scope and selection: The review emphasizes recent literature and selected topics rather than providing a historical or exhaustive overview.The included publications were those considered most relevant or interesting, with citation counts sometimes supporting inclusion decisions.
- Scope and selection: Even the selected topics cannot be covered fully comprehensively because of the size of the literature.The review therefore focuses on publications judged most relevant or interesting.
3. Bibliographic databases
The review compares Web of Science, Scopus, and Google Scholar as sources for citation-impact analysis, emphasizing differences in coverage, access, and data characteristics. Studies generally find broadly similar results, but database choice affects coverage and large-scale usability.
- Database overview: The three principal databases considered for citation analysis are Web of Science, Scopus, and Google Scholar.The review focuses especially on studies of their coverage.
- Database overview: Web of Science and Scopus are subscription-based databases, whereas Google Scholar is a freely available scholarly-literature search engine.Google Scholar provides limited bibliographic metadata and indexes journals, proceedings, books, theses, preprints, and technical reports.
- Google Scholar: Google Scholar’s indexed-document estimates range from about 100 million English-language documents to 160–165 million documents without language restriction.The estimates come from separate studies using different language scopes.
- Web of Science and Scopus: In one comparison, 97% of publications from 2005 covered by Web of Science were also covered by Scopus, making Web of Science nearly a subset of Scopus.The cited comparison concerns matched publications in the two databases.
- Web of Science and Scopus: Scopus generally covers more publications than Web of Science, especially in conference proceedings and several social-science, humanities, engineering, and technology areas.Some studies also find publications covered by Web of Science but not Scopus, so coverage is not identical.
- Comparative results: The three databases produced broadly similar indicators, with Scopus-based indicators correlating most strongly and Web-of-Science-based indicators most weakly with expert judgment, although differences were very small.This comparison concerns agreement with expert judgment rather than database coverage alone.
4. Basic citation impact indicators
The paper organizes citation impact measurement around five basic indicators and distinguishes size-dependent from size-independent measures. It also reviews why average-based indicators can be unstable and limits its scope to output-data indicators.
- Five basic indicators: The five basic indicators are total citations, average citations per publication, highly cited publications, the proportion of highly cited publications, and the h-index.These indicators form the framework for organizing the discussion.
- Indicator properties: Average citation indicators can be strongly influenced by one or a few highly cited publications because citation distributions are highly skewed.The literature therefore sometimes suggests complementing or replacing averages with alternative indicators.
- Five basic indicators: The h-index requires h publications with at least h citations while the remaining publications have no more than h citations; it equals three in the paper’s example.The review notes that many h-index variants exist but does not discuss them in detail.
- Size dependence: Average citations per publication and the proportion of highly cited publications are size-independent, whereas total citations, highly cited publications, and the h-index are size-dependent.Size-independent measures can decrease when publications are added and are commonly used to compare units of different size.
- Scope: The review focuses exclusively on indicators based on publication and citation output data, excluding detailed discussion of measures that also use inputs such as researchers or funding.Input data can support combined measurements of publication productivity and citation impact, but these indicators receive limited attention in the literature reviewed.
5. Selection of publications and citations
The review examines how selecting publications, citations, document types, languages, self-citations, and citation windows shapes citation impact indicators. These choices affect comparability, fairness, and the balance between timely and accurate measurement.
- Selection framework: Citation analyses may select publications by time period, document type, language, journal orientation, or other criteria, and may similarly restrict which citations count.The review explicitly covers excluding publication and citation types from indicator calculations.
- Document type: Different document types are difficult to compare, making exclusions especially important for size-independent indicators such as average citations per publication.The problem is less significant for total citations or the h-index than for average-based indicators.
- Document type: Editorials may receive fewer citations than research articles, so including them can penalize researchers measured by average citations per publication.Excluding editorials is presented as one way to avoid this effect.
- Language and journal orientation: Non-English publications receive fewer citations on average than English publications, and some studies therefore recommend excluding them from comparative size-independent indicators.The literature attributes this difference to the difficulty many researchers have reading non-English publications, while also discussing exclusion of non-international journals.
- Self-citations: At the macro level, author self-citations have a very small effect, and one study reports that each author self-citation yields an additional 3.65 citations from others.These findings support the conclusion that excluding self-citations may not be necessary at country or similar aggregate levels.
- Citation windows: Citation-window evidence is mixed: two or three years may suffice in most fields, but space-research publications are reported to require at least five years.Citation-window choice involves a trade-off between accuracy and timeliness, and short windows are often found relatively insensitive at aggregate levels.
6. Normalization
Normalization is used to compare citation impact across fields, years, and document types by correcting for variables that should not influence evaluation. The reviewed literature addresses ratio-based scores, field definitions, aggregation levels, and cited-side versus citing-side approaches.
- Field differences: Citation density differs greatly across fields, so raw citation counts can reverse apparent impact comparisons, such as 25 biochemistry citations versus 10 mathematics citations.The review notes an approximately order-of-magnitude difference in citation density between biochemistry and mathematics.
- Purpose of normalization: Normalized citation impact indicators correct for field, publication year, and sometimes document-type differences when direct citation comparisons are inappropriate.Other variables, such as author count or publication length, could also be corrected, but their desirability is less clear.
- Limitations: Size-dependent indicators normalized for citation-density differences remain sensitive to differences in publication density across fields.The review identifies this as a limitation of such normalized indicators.
- Ratio-based normalization: A normalized citation score divides a publication’s actual citations by its expected citations, and a research-unit average is obtained by averaging these publication-level scores.The expected value can be based on publications in the same field, year, and document type.
- Ratio-based normalization: In the example research unit, the average normalized citation score is 1.07, indicating that its publications were cited above the expected level on average.The example contains five publications and uses publication-level actual-to-expected citation ratios.
- Alternative approaches: Studies disagree on some normalization choices: Radicchi and Castellano’s approach performs best in one comparison, while aggregation level and cited-side versus citing-side normalization remain contested.Different aggregation levels are described as offering distinct, potentially legitimate viewpoints.
7. Counting methods
The literature compares ways to allocate publication credit among increasingly numerous co-authors. Full counting is simple but can inflate and produce non-additive statistics, while fractional counting often yields lower scores and has no universally preferred interpretation.
- Motivation: Scientific collaboration and author counts are increasing, making it more difficult to allocate publication credit properly to individual authors.Large collaborations can include several hundred authors, particularly in high-energy physics and some biomedical fields.
- Full counting: Full counting gives every author the publication’s full credit, so a five-author paper cited ten times contributes ten citations to each author.This allocates 50 citations overall and therefore counts the same citations multiple times.
- Fractional counting: Fractional counting divides publication credit among authors or participating countries, with allocation rules varying by unit of analysis.Examples include equal shares among authors and country weights based on the authors’ affiliations.
- Empirical comparisons: Empirical comparisons show that fractional counting yields lower citation scores than full counting because internationally co-authored publications receive more citations on average.Fractional weights reduce the contribution of these highly cited multinational publications.
- Interpretation: The literature does not reach consensus because full and fractional counting may measure participation and contribution, respectively.Studies often prefer fractional counting, while full counting is criticized for non-additive country statistics.
- Other counting methods: Author-position methods attempt to reflect unequal contributions, but first- and corresponding-author concepts can be ambiguous and author order is field-dependent.Multiple first or corresponding authors and non-alphabetical ordering conventions limit straightforward interpretation.
8. Citation impact indicators for journals
Journal citation impact indicators include the impact factor and several time-window, distributional, rank-based, and normalized alternatives. The literature also highlights that journal-level indicators may not represent individual publications and require careful interpretation across fields and document types.
- Basic indicators: The impact factor is the best-known journal indicator and essentially represents a journal’s average citations per publication over a defined window.Its classical form uses citations in one year to publications from the previous two years.
- Time windows: The five-year impact factor extends the citation window to address fields where a two-year window is too short.Journal Citation Reports provide both two-year and five-year impact factors.
- Normalization: Field-normalized indicators compare each publication’s actual citations with the expected citations for the same field, year, and document type.The expected value is the average for publications sharing those characteristics.
- Normalization: A rank-based alternative scores a journal by its position within a Web of Science subject category; a journal ranked 10th among 200 receives approximately 0.95.This score indicates that about 95% of journals in the category have lower impact factors.
- Interpretation and use: Journal citation distributions can be highly skewed, so journal-level impact is not necessarily representative of a typical publication.The literature therefore criticizes using journal indicators to evaluate individual publications or their authors.
9. Recommendations for future research
The review concludes that future work should favor theoretically grounded, practically appropriate indicators rather than simply adding more measures. It also recommends exploiting richer contribution and full-text data to improve citation-impact measurement.
- Recommendation 1: The review recommends introducing new citation impact indicators only when they offer clear added value over existing indicators.The large number of proposed indicators and their often overlapping information reduce the need for additional measures.
- Recommendation 2: Future research should strengthen the theoretical foundations of indicators and state their assumptions and construction consequences explicitly.Many field-normalized indicators are tested empirically without a mathematical framework clarifying when normalization is achieved.
- Recommendation 3: Indicator construction should match the needs and expectations of end users, whose interpretation and practical use require more study.Research is needed on how users interpret indicators, draw conclusions, and understand field normalization.
- Recommendation 4: New data sources should be exploited to obtain more sophisticated measurements of citation impact.The review identifies improved author-contribution metadata and full-text availability as important opportunities.
- Recommendation 4: Full-text citation data can support indicators based on citation frequency, citation location, and the surrounding citation context.These measures broaden citation analysis beyond bibliographic metadata and citation counts alone.