Source-linked AI summary
Source normalized indicators of citation impact: An overview of different approaches and an empirical comparison
Ludo Waltman, Nees Jan van Eck
TL;DR
Meaningful cross-field citation comparisons are difficult because fields differ in citation practices and field boundaries are often fuzzy. The paper compares source-normalized indicators with field-classification normalization and finds that MSNCS(1) and MSNCS(3) are preferable to MSNCS(2), while their choice is of limited practical relevance.
Problem
Different fields have substantially different citation practices, while field-classification normalization relies on artificial and potentially overcoarse boundaries.
Method
The paper reviews three source-normalization indicators, compares them empirically with field-classification normalization, and analyzes how journals should be selected for normalization.
Results
MSNCS(1) and MSNCS(3) seem preferable to MSNCS(2), and their strong correlation makes choosing between them of limited practical relevance.
Takeaways & Limitations
Source-normalization approaches may provide a preferable alternative for assessing citation impact across fields, including for journals, universities, research groups, and researchers.
Takeaways & Limitations
The paper calls for additional empirical comparisons across fields and units of analysis, and for criteria to distinguish different journal types.
Abstract
from arXiv · showhide
Different scientific fields have different citation practices. Citation-based bibliometric indicators need to normalize for such differences between fields in order to allow for meaningful between-field comparisons of citation impact. Traditionally, normalization for field differences has usually been done based on a field classification system. In this approach, each publication belongs to one or more fields and the citation impact of a publication is calculated relative to the other publications in the same field. Recently, the idea of source normalization was introduced, which offers an alternative approach to normalize for field differences. In this approach, normalization is done by looking at the referencing behavior of citing publications or citing journals. In this paper, we provide an overview of a number of source normalization approaches and we empirically compare these approaches with a traditional normalization approach based on a field classification system. We also pay attention to the issue of the selection of the journals to be included in a normalization for field differences. Our analysis indicates a number of problems of the traditional classification-system-based normalization approach, suggesting that source normalization approaches may yield more accurate results.
1. Introduction
Meaningful cross-field citation comparisons require normalization because fields differ greatly in citation practices. The paper reviews source normalization as an alternative to field classification, compares approaches empirically, and examines journal selection.
- Motivation: Citation practices differ substantially across fields, making unnormalized citation impact unsuitable for meaningful between-field comparisons.High-citation-density fields such as cell biology may receive more than an order of magnitude more citations per publication than low-density fields such as mathematics.
- Traditional normalization: Traditional normalization assigns publications to fields and compares their citation impact with publications in the same field.The Web of Science system assigns journals to subject categories, so publications inherit the fields of their journals.
- Limitations: Field-classification normalization is limited by fuzzy field boundaries, uncertain field granularity, heterogeneous subfields, and broad-scope journals.These issues make it unclear whether fields defined through journal categories are homogeneous entities.
- Source normalization: Source normalization corrects for field differences using the referencing behavior of citing publications or journals rather than field assignments.The paper uses source normalization as the common term for related citing-side, fractional-counting, and a priori approaches.
- Study design: The study compares three source-normalization approaches with field-classification normalization and also investigates which journals should enter the normalization.The approaches include audience-factor, fractional-citation-counting, and revised-SNIP-based indicators.
- Scope: The normalization approaches can also assess the citation impact of universities, research groups, and individual researchers.
2. Selection of journals
The paper examines whether special journals should be excluded from citation normalization and proposes identifying nationally or regionally oriented journals from their country-address distributions. It uses Kullback-Leibler divergence to quantify this orientation.
- Why journal selection matters: Including low-impact trade, popular, national, or regional journals can distort normalization when their coverage differs across fields.The paper illustrates this problem by contrasting a field containing only regular journals with another also containing special journals.
- Scope and assumption: Excluding special journals could reduce this disadvantage, but accurately distinguishing regular from special scientific journals is difficult.The study therefore does not introduce precise universal criteria and focuses on journals with strong orientation toward one or a few countries.
- Identification method: The study identifies national or regional journals by measuring how strongly their publication-address distributions concentrate on particular countries.For each journal-country pair, it counts country mentions in publication address lists and forms a journal-level country distribution.
- Identification method: A journal’s country distribution is compared with the database-wide country distribution using Kullback-Leibler divergence.Higher divergence indicates stronger national or regional orientation, while a threshold is needed to classify journals.
3. Indicators
The paper evaluates five journal citation indicators: an unnormalized score, a field-classification-based score, and three source-normalized scores that weight citations by citing behavior.
- Unnormalized indicator: The MCS is the average number of citations per publication without normalization for field differences or citation-window length.It resembles the journal impact factor but uses multiple citing years.
- Classification-based indicator: The MNCS compares each publication’s citations with the average citations of publications in its field and publication year.A ratio above or below one indicates citation impact above or below the field-year expectation.
- Source-normalized indicators: Three MSNCS indicators weight each citation according to the referencing behavior of its citing publication or journal.They differ in whether normalization uses citing-journal references, citing-publication references, or both.
- Source-normalized indicators: The MSNCS(1) indicator uses the average number of active references in the citing journal and year, across multiple citing years.Its reference window matches the citation window of the cited publication.
- Source-normalized indicators: The MSNCS(2) indicator instead uses the number of active references in the citing publication, counting only active references within the matched window.This implements a fractional-citation-counting approach adapted to active references.
- Source-normalized indicators: The MSNCS(3) indicator combines citing-publication references with the citing journal’s proportion of publications containing at least one active reference.The added proportion addresses differences between fields in publications lacking active references and makes the indicator similar to revised SNIP.
4. Empirical analysis
The empirical analysis compares source-normalized indicators with MNCS and examines how journal selection affects results. Source normalization appears preferable in several cases, while classification-based normalization and fractional citation counting show important weaknesses.
- 4.1. General statistics: The indicators differ in their average values: MNCS averages exactly one, MSNCS(1) and MSNCS(3) average above one, and MSNCS(2) averages substantially below one.The MSNCS(2) deviation reflects publications without active references, which provide no credits to earlier publications.
- 4.2. Comparison of indicators: The MNCS indicator remains correlated with unnormalized citation impact because overlapping WoS subject categories prevent perfect field normalization.The authors identify this correlation as an artifact of category overlap.
- 4.2. Comparison of indicators: MSNCS(3) can differ substantially from MNCS for individual journals, especially broad-scope journals that fit poorly into field classifications.For multidisciplinary and general medical journals, the authors argue that MSNCS(3) most likely yields more accurate results.
- 4.2. Comparison of indicators: The authors conclude that MSNCS(3) is likely more accurate for at least some journals, but broader generalization requires further analysis.The evidence is presented as journal-specific rather than establishing universal superiority.
- 4.3. Effect of journal selection: Excluding national and regional journals affects MNCS much more than MSNCS(3), with some high-impact journals losing more than half their MNCS value.The corresponding changes for MSNCS(3) are generally small, while results for the other MSNCS indicators are similar.
5. Conclusions
The comparison supports MSNCS(1) and MSNCS(3) over MSNCS(2) and identifies limitations of MNCS-based normalization, especially its sensitivity to journal selection. The authors therefore favor source normalization while calling for broader empirical comparisons, improved field classifications, and clearer journal-selection criteria.
- MSNCS(1) and MSNCS(3) receive the strongest overall support, while MSNCS(2) does not properly normalize for field differences.
- MNCS struggles with broad-scope journals, heterogeneous citation-density fields, and the selection of journals included in an analysis.
- Improved field classification systems defined at the publication level could address limitations of journal-level classification, although such systems are unavailable in many disciplines.
- The authors call for additional empirical comparisons across normalization approaches, fields, and units such as researchers, groups, and universities.
- Future work should develop criteria for distinguishing journal types and exclude journals with very small numbers of active references when using source normalization.