Source-linked AI summary
Characterizing and modeling citation dynamics
Young-Ho Eom, Santo Fortunato
TL;DR
Citation-distribution shape varies with the publication set, discipline, and time window, motivating a systematic analysis of APS citation networks. The paper compares three distributions with KS goodness-of-fit tests and models citation dynamics using preferential attachment with time-dependent attractiveness. The shifted power law best fits the citation distributions, while early citation bursts are reproduced by the proposed model.
Problem
Citation-distribution shape is difficult to characterize universally because it depends on the disciplines, publication years, and observation windows examined.
Method
The paper compares simple power-law, shifted-power-law, and log-normal fits using KS goodness-of-fit tests, then models citation accumulation with time-dependent attractiveness.
Results
The shifted power law is the most reliable fit across observation periods, while the model reproduces empirical citation distributions and citation bursts.
Takeaways & Limitations
Citation dynamics combines broad citation distributions with early-life bursts, requiring a model whose attractiveness changes over time and varies across papers.
Abstract
from arXiv · showhide
Citation distributions are crucial for the analysis and modeling of the activity of scientists. We investigated bibliometric data of papers published in journals of the American Physical Society, searching for the type of function which best describes the observed citation distributions. We used the goodness of fit with Kolmogorov-Smirnov statistics for three classes of functions: log-normal, simple power law and shifted power law. The shifted power law turns out to be the most reliable hypothesis for all citation networks we derived, which correspond to different time spans. We find that citation dynamics is characterized by bursts, usually occurring within a few years since publication of a paper, and the burst size spans several orders of magnitude. We also investigated the microscopic mechanisms for the evolution of citation networks, by proposing a linear preferential attachment with time dependent initial attractiveness. The model successfully reproduces the empirical citation distributions and accounts for the presence of citation bursts as well.
I. INTRODUCTION
Citation distributions depend on the publication set, discipline, and time window, making their functional form difficult to characterize universally. This paper compares competing distributions across APS citation networks and examines citation-accumulation dynamics.
- Citation networks represent papers as vertices and references as directed edges, with citations corresponding to vertex indegree.
- The functional shape of citation distributions remains elusive because results depend on disciplines, publication years, and observation windows.Biology papers are cited more heavily than mathematics papers on average, while older papers have had longer exposure to citations.
- Earlier studies reported different adequate forms, including power laws for highly cited papers, Tsallis distributions for whole distributions, and log-normal fits.
- The paper analyzes APS citation networks over several time windows using KS goodness-of-fit tests for simple power-law, shifted-power-law, and log-normal models.It also investigates rapid citation accretions, or bursts, and proposes a citation-attractiveness model incorporating time dependence.
A. The distribution of cites
The study selects citation-distribution models by comparing their empirical and fitted cumulative distributions with KS-based goodness-of-fit testing. Across APS observation periods, the shifted power law provides the most reliable fit.
- The KS statistic D is the maximum distance between empirical and fitted cumulative distribution functions.The fit is evaluated in the region k_in ≥ k_min.
- The best model is selected by searching parameters for the least empirical KS distance D.
- Synthetic datasets calibrated to the best-fit curve provide p-values for assessing whether each fitted model is plausible.The p-value is the fraction of synthetic KS statistics larger than the empirical value; 1000 synthetic distributions were used.
- The shifted power law gives significant p-values above 0.2 across all observation periods, unlike the other tested distributions.Simple power laws fit well mainly in the right tail, while log-normal fits deteriorate after 1970 and disagree in the tail.
- The shifted-power-law exponent γ decreases from 5.6 in 1950 to 3.1 in 2008.
- The authors conclude that the shifted power law best fits the APS citation data.
B. The distribution of citation bursts
Citation bursts have broad size distributions and are concentrated early in a paper’s life. The analysis uses one-year observation windows and compares burst statistics across publication-age restrictions.
- Figure 1 and Table I evaluate log-normal, simple power-law, and shifted-power-law fits using Kolmogorov-Smirnov goodness-of-fit statistics.The figure and table identify the empirical distributions and the three candidate model classes.
- Burst-size distributions span several orders of magnitude across citation datasets.The distributions of relative citation-rate changes are broad and heavy-tailed.
- More than 90% of large bursts (∆k/k > 3.0) occur within the first 4 years since publication.When papers older than 5 or 10 years are analyzed, the tail of the burst-size distribution disappears.
C. Preferential attachment and age-dependent attractiveness
The paper tests linear preferential attachment with vertex attractiveness to model citation-network growth. The attractiveness is important early in a paper’s life but loses influence as papers age.
- Preferential attachment: Linear preferential attachment assigns a citation probability proportional to a paper’s indegree plus its attractiveness.
- Preferential attachment: The cumulative-kernel analysis estimates the average attractiveness by aggregating edge-acquisition probabilities across indegree classes.The time window must preserve network structure while providing sufficient citation statistics.
- Preferential attachment: For the 2007–2008 network, the kernel is compatible with linear preferential attachment with average attractiveness ⟨A⟩=7.0 over a large range.The tail slope is close to 2, although the final part of the tail is missed.
- Age-dependent attractiveness: Across datasets from 1950 to 2008, the average attractiveness is 7.1 and dominates preferential attachment at low indegrees.Attractiveness is particularly important for old papers in the early ages of the network.
- Age-dependent attractiveness: For papers older than 5 or 10 years, the kernel becomes initially quadratic in indegree, indicating that attractiveness no longer affects citation dynamics.The paper therefore links attractiveness primarily to the first few years after publication.
D. The model
The model combines linear preferential attachment with paper-specific attractiveness that decays exponentially over time. With heterogeneous initial attractiveness, it reproduces empirical citation and burst-size distributions across citation networks of different ages.
- Model mechanism: Citation probability follows linear preferential attachment weighted by each target paper’s time-dependent attractiveness.Constant equal attractiveness would recover standard linear preferential attachment; the model instead lets attractiveness decay exponentially.
- Model mechanism: A paper’s attractiveness decays exponentially, with τ defining the timescale after which it loses considerable importance for citation dynamics.A0 is initial attractiveness and t0 is the paper’s first appearance in the network.
- Model mechanism: Heterogeneous power-law initial attractiveness accounts for broad citation-burst sizes during papers’ early life.The model links the relevance of initial attractiveness to early citation bursts and their broad size distribution.
- Simulation setup: α = 2.5 and τ = 1 year were used in simulations, with A0 bounded by Amin ≤ A0 < 0.002N(t).The reported bounds vary for Amin: 25.0 for most years and 14.5 for 1950.
- Empirical agreement: The model reproduces empirical citation distributions from 1950 through 2008 and accurately reproduces burst-size distributions for one-year and longer observation windows.For five- and ten-year burst windows, the model accurately describes the empirical curve tails.
III. DISCUSSION
The study finds that shifted power laws best describe APS citation distributions, while citation dynamics shows early-life bursts. Its preferential-attachment model with decaying heterogeneous attractiveness reproduces citation and burst-size distributions across scientific ages.
- Citation distributions: Shifted power laws are the best-supported citation-distribution ansatz across APS networks spanning different time periods.They outperform simple power laws and log-normals under Kolmogorov-Smirnov goodness-of-fit tests.
- Citation dynamics: Citation bursts typically occur during the early life of papers, with burst sizes spanning the dynamics studied across citation networks.The discussion connects these bursts to analogous popularity dynamics in Wikipedia and the Web.
- Model interpretation: Traditional preferential attachment produces smooth citation accumulation and therefore does not account for observed citation bursts.The paper introduces time decay and heterogeneous attractiveness as two additional model features.
- Model interpretation: The proposed model accurately describes citation and burst-size distributions across scientific ages and remains fairly robust to the burst-observation window.The robustness claim concerns the choice of observation window for detecting bursts.
IV. MATERIALS AND METHODS
The citation database covers APS journal papers from 1893 through 2008, excluding Reviews of Modern Physics. It contains 414,977 papers and 3,992,736 citations at the end of 2008.
- Dataset: The database includes papers published in APS journals from 1893 to 2008, excluding Reviews of Modern Physics.The considered journals include Physical Review titles and related APS series.
- Dataset: 414,977 papers and 3,992,736 citations are included at the end of 2008.These totals describe the assembled citation database.