Source-linked AI summary
Power laws in citation distributions: Evidence from Scopus
Michal Brzezinski
TL;DR
The paper addresses empirical detection of power-law behavior in citation distributions amid controversy over which statistical model fits best. Using a large Scopus dataset and rigorous statistical comparisons, it finds that plausible power-law exponents range from 3.24 to 4.69 and that power-law distributions usually account for less than 1% of articles.
Problem
The paper examines how to empirically detect power-law behavior in citation distributions amid disagreement over which model fits citation data best.
Method
The study uses a very large, previously unused dataset, rigorously compares power-law and alternative models, and applies goodness-of-fit and model-selection tests.
Results
Power-law exponents have a plausible range of 3.24 to 4.69, while power-law distributions usually account for less than 1% of articles published in a field.
Takeaways & Limitations
The findings support treating power-law behavior in citation distributions as limited to a small fraction of published articles, with exponents in the reported range.
Takeaways & Limitations
The authors state that the relevant issue should be further studied using appropriate goodness-of-fit tests.
Abstract
from arXiv · showhide
Modeling distributions of citations to scientific papers is crucial for understanding how science develops. However, there is a considerable empirical controversy on which statistical model fits the citation distributions best. This paper is concerned with rigorous empirical detection of power-law behaviour in the distribution of citations received by the most highly cited scientific papers. We have used a large, novel data set on citations to scientific papers published between 1998 and 2002 drawn from Scopus. The power-law model is compared with a number of alternative models using a likelihood ratio test. We have found that the power-law hypothesis is rejected for around half of the Scopus fields of science. For these fields of science, the Yule, power-law with exponential cut-off and log-normal distributions seem to fit the data better than the pure power-law model. On the other hand, when the power-law hypothesis is not rejected, it is usually empirically indistinguishable from most of the alternative models. The pure power-law model seems to be the best model only for the most highly cited papers in "Physics and Astronomy". Overall, our results seem to support theories implying that the most highly cited scientific papers follow the Yule, power-law with exponential cut-off or log-normal distribution. Our findings suggest also that power laws in citation distributions, when present, account only for a very small fraction of the published papers (less than 1% for most of science fields) and that the power-law scaling parameter (exponent) is substantially higher (from around 3.2 to around 4.7) than found in the older literature.
Introduction
The paper addresses whether power laws describe citation distributions among highly cited scientific papers, using large Scopus data and rigorous comparisons with alternative models.
- Motivation: Power-law citation models are motivated by broader efforts to understand heavy-tailed scientific and social phenomena.Related examples include author productivity, word occurrence, citations, and network structure; equivalent laws are called Pareto’s law in economics and Zipf’s law in linguistics.
- Research question: Power-law behavior in citation distributions is not universal across scientific fields.Prior WoS studies found that power laws could not be rejected in 17 of 22 fields and 140 of 219 sub-fields.
- Research question: The paper tests whether a power-law model best describes the observed distribution of highly cited papers.It uses a statistical toolbox for detecting power-law behavior and addresses model-selection concerns in prior work.
- Data and coverage: 2.2 million articles published between 1998 and 2002 across 27 Scopus subject areas form the paper’s main dataset.Scopus indexes about 70% more sources than Web of Science, providing broader citation coverage; its most cited paper received 5187 citations.
- Methods and contribution: Earlier studies often used small datasets, single alternatives, or visual model-selection methods that were unsuitable for rigorous comparison.The Scopus sample is larger for highly cited papers than recent large WoS samples.
- Methods and contribution: The paper systematically compares the pure power-law model with log-normal, exponential, stretched-exponential, Tsallis, Yule, and truncated power-law alternatives.Model comparisons use a likelihood ratio test rather than relying only on visual inspection.
Materials and Methods
The study fits discrete power-law models to citation counts above an estimated threshold, tests their goodness of fit, and compares them with alternative distributions using likelihood ratios. Citation data come from Scopus, whose coverage is judged useful for analyzing right tails but has documented coverage limitations.
- Power-law model: Citation counts are modeled as a discrete power law above a minimum value x0, with α as the scaling exponent.The model uses the Hurwitz zeta function for normalization.
- Parameter estimation: x0 is estimated by selecting the threshold that minimizes the Kolmogorov-Smirnov statistic between observed and fitted cumulative distributions.The exponent is estimated by maximum likelihood, and parameter uncertainty uses 1,000 bootstrap replications.
- Goodness-of-fit testing: Goodness of fit is assessed by semi-parametric bootstrap simulation, rejecting the power-law model when the estimated p-value is below 0.1.Synthetic data retain the empirical distribution below the estimated threshold and follow the fitted power law above it.
- Model comparison: Likelihood ratio tests compare the power law with exponential, stretched exponential, log-normal, Yule, power-law cutoff, and Tsallis alternatives.A model is favored only when the normalized likelihood-ratio comparison is significantly different from zero; otherwise, the test cannot choose between models.
- Data: The dataset contains Scopus citation records for articles, with the analysis focused on the right tails of citation distributions.Scopus is considered useful for this purpose because its highest-citation coverage is satisfactory, although some fields contain only partial distributions.
Results and Discussion
Across Scopus fields, pure power laws receive weak and uneven support: they are rejected for roughly half the fields, while alternatives often fit better or remain statistically indistinguishable. Where plausible, power laws cover a small share of papers and have exponents around 3.24–4.69.
- The power-law hypothesis cannot be rejected for 14 Scopus science fields, including Physics and Astronomy, Chemistry, and Multidisciplinary.
- 13 Scopus fields reject the power-law model, spanning humanities, social sciences, formal sciences, life sciences, Earth and Planetary Sciences, and Engineering.
- For most distributions, faster-than-power-law tail decay suggests using lighter-tailed log-normal or power-law-with-exponential-cutoff models.
- 3.24–4.69 is the estimated exponent range for the 14 Scopus fields where the power law remains plausible, substantially above earlier estimates around 2.3–3.
- Less than 1% of articles usually belong to the power-law region, with Chemistry at 2% and Multidisciplinary at 2.8%.
- The Yule and power-law-with-exponential-cutoff models significantly outperform the pure power law in nearly all rejected fields except Veterinary, while log-normal performs better in 10 fields.
- Among fields passing goodness-of-fit, pure power law is favored only for Physics and Astronomy; elsewhere, alternatives often have higher likelihoods but are usually statistically indistinguishable.
- Overall, the evidence favors Yule, power-law-with-exponential-cutoff, or log-normal distributions over pure power laws, although alternative fits require further goodness-of-fit testing.
Conclusions
Using a large Scopus dataset covering papers published in 1998–2002, the paper tests power-law behavior in citation-distribution right tails. Power laws are rejected in around half of fields, while plausible cases generally cannot distinguish them from alternatives and cover less than 1% of papers.
- The study tests power-law behavior in citation-distribution right tails using a large Scopus dataset of papers published between 1998 and 2002.
- Around half of Scopus fields reject the power-law hypothesis.
- In the remaining fields, power laws are plausible, but differences from alternative models are usually statistically insignificant.
- Less than 1% of published papers follow plausible power laws in most science fields.
- The study confirms that citation-distribution power-law exponents are substantially higher than reported in older literature.
Figure Legends
Figure 1 compares empirical citation-distribution tails with their best power-law fits for distributions that failed the goodness-of-fit test.
- Blue circles show complementary cumulative distribution functions, while dashed black lines show the best power-law fits.The plotted distributions are from Scopus papers published in 1998–2002 using a 5-year citation window.
Tables
The tables define alternative discrete distributions, describe the citation data, report power-law fits, and present model-selection tests.
- Table 1 defines the alternative discrete distributions used for comparison with the power-law model.The alternatives were selected following Clauset et al.
- The distributions are normalized so total probability over [x0, +∞] equals 1; the discrete log-normal is approximated by rounding continuous values.The Tsallis distribution uses a parametrization considered by Shalizi.
- Table 3 reports discrete power-law fits to the citation datasets.
- Table 4 presents model-selection tests, where positive likelihood-ratio values indicate that the power-law model is favored.Standard errors are reported in parentheses, and the second column gives the p-value for the power-law hypothesis.