Source-linked AI summary
Universality of citation distributions: towards an objective measure of scientific impact
Filippo Radicchi, Santo Fortunato, Claudio Castellano
TL;DR
Citation analysis needs a fairer way to compare publications across disciplines and years because raw citation distributions vary substantially. The paper empirically tests the relative indicator cf = c/c0, finds a universal rescaled citation distribution, and uses it to generalize the h-index across fields.
Problem
Raw citation counts vary across fields, while citation distributions are skewed even within disciplines, making fair cross-field evaluation difficult.
Method
The paper empirically compares single-publication citation distributions using cf = c/c0, where c0 is the same-year field average, and extends the indicator to a generalized h-index.
Results
Citation distributions from different disciplines and years rescale onto a universal curve, supporting cf as an unbiased indicator for comparing publication impact.
Takeaways & Limitations
Relative citation performance should be used instead of bare citation counts for cross-field comparisons, and cf also supports comparing scientists through a generalized h-index.
Abstract
from arXiv · showhide
We study the distributions of citations received by a single publication within several disciplines, spanning broad areas of science. We show that the probability that an article is cited $c$ times has large variations between different disciplines, but all distributions are rescaled on a universal curve when the relative indicator $c_f=c/c_0$ is considered, where $c_0$ is the average number of citations per article for the discipline. In addition we show that the same universal behavior occurs when citation distributions of articles published in the same field, but in different years, are compared. These findings provide a strong validation of $c_f$ as an unbiased indicator for citation performance across disciplines and years. Based on this indicator, we introduce a generalization of the h-index suitable for comparing scientists working in different fields.
I. INTRODUCTION
Citation counts vary systematically across scientific fields, making raw counts unsuitable for fair comparison. The paper examines whether normalizing by field- and year-specific average citations provides an unbiased relative indicator despite skewed citation distributions.
- Motivation: Field variation is a major source of bias in evaluating scientific performance because disciplines receive systematically different citation frequencies.The paper attributes this variation partly to differences in the number of references per article and cross-discipline citation patterns.
- Related work: Comparing bare citation counts is inappropriate, motivating normalization against a carefully chosen reference standard.Existing approaches include within-field rankings and ratios of citations to field- or journal-based averages.
- Research question: Citation distributions are highly skewed even within disciplines, raising doubts about whether the average citation count adequately characterizes them.The paper explicitly addresses this concern through an empirical analysis of article-citation distributions.
- Approach: The study evaluates single publications using the average citations c0 received by articles in the same scientific field and publication year.The reference set consists of papers in journals classified within the same Journal of Citation Reports scientific category.
- Contribution: The relative indicator cf = c/c0 is proposed to account for variation across fields and years, with the paper testing whether its distribution is shared across disciplines.The stated goal is to validate cf as an unbiased relative indicator of scientific impact.
II. VARIABILITY OF CITATION STATISTICS IN DIFFERENT DISCIPLINES
Article-citation distributions differ strongly across disciplines, so identical raw citation counts do not represent equivalent performance. Rescaling citations by the field average produces a common distribution across the disciplines considered.
- Disciplinary variability: Citation distributions strongly depend on discipline: a 100-citation publication is approximately 50 times more common in Developmental Biology than in Aerospace Engineering.The comparison uses articles published in 1999 across journals classified in several Journal of Citation Reports disciplines.
- Normalization: The rescaled distribution c0P(c, c0) of cf = c/c0 collapses the citation curves for the disciplines considered onto a universal curve.The relative indicator divides citations by the discipline’s average citation count per article.
- Implications: A 100-citation article in Developmental Biology cannot be judged more successful than an Aerospace Engineering article using citation count alone.The paper presents this as a direct implication of the observed field-dependent citation frequencies.
III. DISTRIBUTION OF THE RELATIVE INDICATOR cf
Rescaling citation counts by the field average c0 produces a nearly universal relative-performance distribution across disciplines and supports fairer cross-field rankings. The same relative indicator also remains comparable across publication years.
- Relative indicator: cf = c/c0 collapses large discipline-specific citation differences into a universal curve, with c0 defined as the field’s average citations per article.The reference standard uses articles from the same field and publication year.
- Relative indicator: A single lognormal fit across all categories gives σ2 = 1.3.The fitted parameter values are generally compatible within two standard deviations, with Anesthesiology within three.
- Cross-field ranking: For z = 20%, cf-based ranking has σz = 1.15%, versus 12.37% for c-based ranking and 1.09% under unbiased ranking.These results indicate close agreement between cf-based rankings and the theoretical unbiased benchmark.
- Implications: The universal relative-indicator distribution supports comparing publication impact across fields, while its cross-year stability supports comparisons across publication years.The indicator is interpreted as a publication’s citation performance relative to its same-year, same-field average.
- Across years: The rescaled citation distribution remains conspicuously the same across different publication years within a discipline, even though c0 increases for older articles.This extends the comparison from cross-disciplinary normalization to longitudinal comparisons.
IV. A GENERALIZED H-INDEX
The generalized h-index combines citation normalization by cf with publication-rate normalization by N0, enabling comparisons among scientists in different disciplines.
- Motivation: The h-index is difficult to compare across disciplines because citation patterns and publication rates vary by field.The paper addresses both sources of variation by ranking articles with cf and scaling ranks by N0.
- Temporal stability: Citation distributions for articles published in different years retain the same universal scaling after rescaling by c0.Figure 4 compares Hematology, Neuroimaging, and Physics, Nuclear for 1990, 1999, and 2004; the dashed curve is a lognormal fit with σ^2 = 1.3.
- Publication-rate normalization: Publication-count distributions collapse onto a universal curve when authors’ annual output N is divided by the disciplinary average N0.The rescaled distribution follows a power law over almost two decades, with exponent δ = 3.5(5).
- Definition of hf: The generalized index hf orders an author’s articles by cf and compares each relative citation value with the reduced rank r/N0.hf is the last value of r/N0 for which cf exceeds r/N0.
- Example: An author with cf values 4.1, 2.8, 2.2, 1.6, 0.8, and 0.4 and N0 = 2.0 has hf = 1.5.The third article satisfies 1.5 < 2.2, whereas the fourth does not satisfy 2.0 < 1.6.
V. CONCLUSIONS
The paper concludes that cf places citation distributions from different disciplines and years on a common curve, supporting fairer article comparisons. It also extends normalization to publication rates while identifying author number and the origin of universality as open issues.
- Conclusions: Citation distributions from different disciplines collapse onto the same universal curve when citations are expressed as cf.The universal curve is reported to remain remarkably stable over the years.
- Implications: Relative indicators support fairer comparisons of article impact across disciplines and years, but higher cf does not necessarily mean greater importance.Field-related factors may affect broader scientific or societal importance.
- Scope and applications: The analysis concerns single publications, so evaluations of individuals or research groups must also account for additional complexities.For mean citation performance, the paper recommends averaging cf rather than bare citation counts; for authors, it defines a generalized h-index incorporating publication rates.
- Open limitations: The study addresses two major sources of citation-comparison bias but leaves author number as a potential additional source.The paper proposes investigating whether citations per author produce a universal distribution.
- Future work: Explaining the origin and precise functional form of the observed universality remains a theoretical goal for future work.The paper points to mechanisms underlying citation distributions as an open research direction.
VI. METHODS
The empirical analysis uses Thomson Scientific’s Web of Science to count citations and classify journals into scientific categories. Articles in journals within each category are treated as belonging to that category.
- Data source: Citations are counted as references to an article in more recent published articles.The source database is Thomson Scientific’s Web of Science.
- Field classification: Web of Science divides scientific journals into 172 categories ranging from Acoustics to Zoology.Each category contains a list of journals used to define the corresponding field.
- Field classification: Articles published in journals listed within a category are treated as part of that category.This category assignment provides the field grouping used in the empirical analysis.