Source-linked AI summary
A systematic empirical comparison of different approaches for normalizing citation impact indicators
Ludo Waltman, Nees Jan van Eck
TL;DR
The paper asks how citation-based indicators can be normalized for fair comparisons across fields and years. It systematically compares one field-classification approach with three source-normalization approaches using a carefully selected publication set and algorithmically constructed classification systems. Two source-normalization approaches generally outperform the classification-based approach, while the fractional citation-counting approach does not perform well.
Problem
The paper addresses how citation-based indicators can be normalized to support fair comparisons across scientific fields and years.
Method
The study performs a systematic large-scale comparison of one field-classification approach and three source-normalization approaches using selected WoS core-journal publications and independent algorithmic classification systems.
Results
Two source-normalization approaches generally perform better than the classification-system-based approach, especially at higher levels of granularity.
Takeaways & Limitations
The analysis supports using a source-normalization approach, except at very low levels of granularity.
Takeaways & Limitations
The analysis is limited because any classification system can be subject to classification-related limitations.
Abstract
from arXiv · showhide
We address the question how citation-based bibliometric indicators can best be normalized to ensure fair comparisons between publications from different scientific fields and different years. In a systematic large-scale empirical analysis, we compare a traditional normalization approach based on a field classification system with three source normalization approaches. We pay special attention to the selection of the publications included in the analysis. Publications in national scientific journals, popular scientific magazines, and trade magazines are not included. Unlike earlier studies, we use algorithmically constructed classification systems to evaluate the different normalization approaches. Our analysis shows that a source normalization approach based on the recently introduced idea of fractional citation counting does not perform well. Two other source normalization approaches generally outperform the classification-system-based normalization approach that we study. Our analysis therefore offers considerable support for the use of source-normalized bibliometric indicators.
1. Introduction
The paper examines how citation-based indicators should be normalized to enable fair comparisons across fields and identifies source-normalization approaches as a promising alternative to field classification.
- Motivation: Citation practices differ across scientific fields, making raw citation counts unsuitable for fair between-field comparisons.Fields vary in reference-list length, reference age, and average citations received.
- Research gap: Earlier systematic comparisons are scarce and face methodological challenges, with some reporting greater accuracy for classification-system-based approaches.The authors characterize earlier results as methodologically limited and therefore insufficient for resolving which approach performs best.
- Research question and approach: The paper compares field-classification normalization with three source-normalization approaches in a systematic large-scale empirical analysis.The comparison includes one classification-system-based approach and three source-normalization approaches.
- Methodological design: The study excludes publications likely to distort comparisons, including national scientific journals, popular scientific magazines, and trade magazines.The analysis focuses on selected sources intended to represent the international scientific literature covered by Web of Science.
- Methodological design: The evaluation uses classification systems different from the implementation system and compares performance at broad-discipline and subfield levels.Algorithmically constructed systems assign publications to fields based on citation patterns, while different granularity levels can yield different performance.
- Evaluation scope: The study reports that some normalization approaches perform well at one granularity level but not at another.This motivates evaluating methods across both broad scientific disciplines and smaller scientific subfields.
2. Data
The analysis uses Web of Science data restricted to selected core journals, with publications and citations defined within that selected set and field systems constructed at different granularities.
- Data source and selection: The dataset combines Web of Science indexes and focuses on selected publications from WoS core journals.The selection is intended to restrict the analysis to international scientific literature covered by Web of Science.
- Data source and selection: The study excludes special source types because their presence may distort normalization for differences in citation practices between fields.Examples include national or regional journals, trade magazines, business magazines, and popular scientific magazines.
- Data source and selection: Of 9.79 million Web of Science article and review records, 8.20 million publications from 2003–2011 are included in the analysis.The selected publication set is defined after applying the core-journal selection procedure.
- Citation data: Citations and references are counted only when both citing and cited publications belong to the selected WoS core-journal set.Citations from non-selected publications and references to non-selected publications therefore play no role.
- Citation data: Citation impact is calculated for 3.86 million publications from 2007–2010 using citations counted through the end of 2011.The analysis contains 26.22 million citations, with a relatively short two- to five-year citation window.
3. Normalization approaches
The paper contrasts field-based normalization with source normalization, which weights citations according to citing behavior while also correcting for publication age and field differences.
- Age correction: Because older publications have longer citation windows than newer publications, the normalization approaches also need to correct for publication age.The analysis compares publications from 2007–2010 while counting citations through the end of 2011.
- Classification-system-based normalization: The classification-system-based approach computes a normalized citation score by comparing actual citations with the field-year average.The expected citation count is based on publications in the same field and year.
- Classification-system-based normalization: An NCS above or below one indicates citation impact above or below the expected level for the publication’s field and year.Averaging publication-level NCS values produces the mean normalized citation score indicator.
- Classification-system-based normalization: The classification-based method uses Web of Science subject categories to determine expected citation counts, despite their limitations.Publications assigned to multiple subject categories use the harmonic average of the category-specific expected citation counts.
- Source normalization: The three source-normalization approaches weight each received citation according to the referencing behavior of the citing publication or journal.They differ in the precise rule used to determine each citation’s weight.
- Source normalization: An active reference is a reference within a specified window that points to a publication in a WoS core journal.References to sources outside the WoS database or to non-core journals are not counted as active references.
4. Results
The results compare four normalization approaches using Web of Science subject categories and algorithmically constructed classification systems. The latter evaluation is presented as fairer and generally favors source normalization, while SNCS(2) performs poorly because it does not properly correct for publication age.
- Evaluation design: Using algorithmically constructed classification systems A, B, and C is argued to provide a fairer comparison of the normalization approaches.The Web of Science subject-category evaluation is described as potentially biased in favor of NCS.
- Results based on Web of Science subject categories: The average SNCS(2) value in 2010 is more than 30% lower than in 2007, indicating inadequate correction for publication age.Recent publications therefore have a significant disadvantage compared with older ones.
- Results based on Web of Science subject categories: SNCS(1) and SNCS(3) have average yearly values 5% to 10% above one, whereas SNCS(2) is consistently outperformed by them.SNCS(1) and SNCS(3) yield the same average values per year with a small decreasing trend over time.
- Results based on Web of Science subject categories: The Web of Science evaluation favors NCS, but this advantage is attributed to bias from using the same classification system for implementation and evaluation.This can prevent weaknesses in the classification system from being detected and overestimate NCS performance.
5. Conclusions
The analysis supports source normalization for fairer citation-impact comparisons, while showing that performance depends on the normalization approach and granularity. It also identifies methodological biases and scope limitations affecting classification-based evaluations.
- Methodological innovations: Using the same classification system to implement and evaluate normalization produces significantly biased results.The authors therefore distinguish the classification system used in implementation from the one used in evaluation.
- Approach: The analysis compared one field-classification normalization approach with three source-normalization approaches across different levels of granularity.The study also excluded special publication types and used algorithmically constructed classification systems for evaluation.
- Findings: The fractional citation-counting approach does not provide a completely satisfactory normalization and fails to properly correct for publication age.The authors advise against source-normalization approaches following this idea.
- Findings: The other two source-normalization approaches generally outperform the WoS subject-category approach, especially at higher granularity levels.The revised SNIP approach is more accurate than the audience-factor approach except at very low granularity, such as broad-discipline comparisons.
- Limitations: Classification-based normalization remains subject to artificial field boundaries, uncertain granularity choices, and possible instability when algorithmically constructed systems are updated.The evaluation systems also assign each publication to exactly one research area, leaving no room for multidisciplinary publications.
- Limitations: The analysis assumes that successful normalization should make citation distributions across fields coincide as much as possible, although alternative correction criteria may exist.The authors note that their algorithmically constructed evaluation systems have limitations similar to other classification systems.
Appendix A: Selection of Web of Science core journals
Core journals are selected from Web of Science publications through three filtering steps addressing document type, international orientation, and connection to recent indexed literature.
- Selection criteria: The procedure starts with Web of Science publications classified as articles or reviews from 1999–2011.
- International orientation: Journals are evaluated for international orientation using publication and citing-publication country distributions compared through Kullback–Leibler divergence.Higher divergence indicates stronger concentration in one or a few countries.
- International orientation: Journals with divergence values below 1.3260 are treated as international, while the remaining journals are excluded as core journals.The threshold is acknowledged to involve some arbitrariness.
- Connection to indexed literature: The final step requires at least 50% of a journal’s publications to cite a recent publication in a Web of Science core journal.The cited publication must be no more than four years older, and exclusions are iterated until the journal set stabilizes.
- Connection to indexed literature: The resulting core-journal set is iteratively reduced until all remaining journals satisfy the 50% citation requirement.
Appendix B: Algorithmic construction of field classification
Three field classification systems are constructed algorithmically from the direct-citation network, differing in granularity and assigning publications to research areas.
- Classification granularity: Systems A, B, and C contain 21, 161, and 1,334 research areas, respectively, representing increasing levels of detail.
- Network construction: The classification systems are built from direct citation relations among 8.20 million publications, with citation direction ignored.The network contains 80.56 million direct citation relations.
- Research-area assignment: A normalization procedure and clustering technique assign publications to research areas while accounting for differences in citation practices between disciplines.
- Coverage: Publications outside the largest direct-citation component are added to the research area with which they are most strongly bibliographically coupled.
- Coverage: Each system ultimately includes 8,117,743 publications, leaving 82,084 publications, or 1.0%, unclassified.The excluded publications have neither direct citation nor bibliographic-coupling relations to included publications.
Appendix C: Decomposition of citation inequality
Citation inequality is decomposed into within-quantile, quantile-distribution, and citation-practice components to evaluate normalization across fields and publication ages.
- Decomposition: The decomposition separates overall citation inequality into within-quantile inequality, inequality in normalized citation values, and inequality from citation-practice differences.
- IDCP measure: IDCP is a weighted average of quantile-specific inequality indices, weighted by the total citation score of each quantile interval.
- IDCP measure: Lower IDCP values indicate better correction for differences in citation practices between fields and differences in publication age.
- Evaluation: The analysis calculates I, W, S, and IDCP for four normalization approaches and an unnormalized classification-system approach across three classification systems.
- Evaluation: Differences between approaches’ IDCP values are relatively small because the highest quantile intervals receive substantial weight and show limited differences across approaches.