Source-linked AI summary

Diffusion of scientific credits and the ranking of scientists

Filippo Radicchi, Santo Fortunato, Benjamin Markines, Alessandro Vespignani

arXiv:0907.1050v2physics.soc-phcs.DLphysics.data-an

TL;DR

The paper addresses whether scientist rankings should account for the non-local diffusion of scientific credit rather than rely only on citation counts. It constructs weighted author citation networks from the Physical Review archive, applies the SARA diffusion algorithm, and compares it with local metrics using major Physics prizes as a benchmark. SARA generally places most prize-winning scientists near the top of the ranking, while its scope is limited by the Physical Review dataset and author-name disambiguation issues.

  • Problem

    Standard citation-based indicators may not capture scientists’ actual merit because citations differ in value according to the citing scientist and credit diffusion is non-local.

  • Method

    SARA projects normalized paper citations into a weighted author network and combines biased random walks with random credit redistribution to diffuse scientific credit.

  • Results

    SARA generally identifies major-prize winners by placing most of them in top ranking positions.

  • Takeaways & Limitations

    The ranking captures non-local effects of scientific-credit spreading and provides an alternative to citation counts and related local metrics.

  • Takeaways & Limitations

    The algorithm uses only the Physical Review dataset, which can undervalue authors whose impact extends into other disciplines.

Abstract

from arXiv · show

Recently, the abundance of digital data enabled the implementation of graph based ranking algorithms that provide system level analysis for ranking publications and authors. Here we take advantage of the entire Physical Review publication archive (1893-2006) to construct authors' networks where weighted edges, as measured from opportunely normalized citation counts, define a proxy for the mechanism of scientific credit transfer. On this network we define a ranking method based on a diffusion algorithm that mimics the spreading of scientific credits on the network. We compare the results obtained with our algorithm with those obtained by local measures such as the citation count and provide a statistical analysis of the assignment of major career awards in the area of Physics. A web site where the algorithm is made available to perform customized rank analysis can be found at the address http://www.physauthorsrank.org

I. INTRODUCTION

Digital bibliographic data support graph-based rankings of scientific impact, but standard citation-based indicators may not capture how scientific credit flows among authors. The paper introduces a diffusion-based approach using the Physical Review archive and compares it with existing ranking schemes and prize-based benchmarks.

  • Digitalization has enabled system-level graph analyses and quantitative rankings of journals, papers, people, and disciplines.
  • Common indicators such as citation counts and h-indexes are criticized for not fully representing a scientist’s actual merit.
  • Citations have different values depending on the citing scientist, making scientific credit diffusion a non-local process.
  • The paper applies a diffusion algorithm to 407 236 Physical Review papers published between 1893 and 2006.
  • SARA is compared with Citation Count and Balanced Citation Count, then evaluated against major Physics prize winners.

II. DESCRIPTION OF THE DATASET

The dataset comprises Physical Review publications and their references, with author, bibliographic, and citation information extracted from publisher-provided XML files. The analysis uses a defined reference set that includes citations to papers inside and outside the Physical Review journals.

  • The database contains 407 236 papers published from 1893 to 2006 across nine Physical Review journals.
  • Publisher-provided XML files supply authors, dates, journals, volumes, pages, references, PACS numbers, and additional information.
  • The paper citation network is constructed from 9 359 556 references, including 3 866 471 internal Physical Review references.
  • After excluding “First author et al.” references and papers by authors without Physical Review publications, 8 783 994 references remain.
  • The analysis uses all 8 783 994 retained references, including references to papers published outside the Physical Review journals.

III. CONSTRUCTION OF THE WEIGHTED AUTHOR CITATION NETWORK

The weighted author citation network is obtained by projecting each paper citation onto all citing–cited author pairs and distributing the citation weight across those pairs. This projection introduces author-name disambiguation as a central data-quality issue.

  • A paper citation from n authors to an m-author paper creates n · m directed author connections.
  • The network is a projection of the paper citation network, using reference links to construct weighted directed author links.
  • Figure 1 illustrates that two citing authors assign weights 1/2 to a one-author target and 1/4 to each author of a two-author target.
  • Each citing–cited author connection receives weight 1/(n · m), and weights are summed across all references between the same authors.
  • Author projection creates name-disambiguation problems because common names and incomplete author names can identify multiple people.

A. Dynamical Representation of the Weighted Author Citation Network

To preserve temporal information and account for rising citation rates, the authors divide chronologically ordered references into overlapping slices with equal reference counts. The resulting intervals represent unequal periods of calendar time.

  • A single network would mix very old and new citations and lose longitudinal information as citation rates increase over time.
  • The reference list is sorted by the publication date of each citing paper before slicing.
  • Figure 2 displays only links above a threshold; connection width represents weight and node size represents total incident-link weight.
  • References are divided into MI homogeneous intervals, each containing the same number MR of references.
  • Adjacent intervals overlap by sharing references at their boundaries to avoid abrupt changes.
  • With MR = 488 000, the first interval spans 1893–1966 while the last represents 2006, and intervals are not shorter than one year.

B. Properties of the Weighted Author Citation Network

The weighted author citation networks change over time as the physics community grows. Rescaled indegree distributions show common behavior across periods, whereas instrength distributions do not collapse onto a universal curve.

  • The network’s author, indegree, and instrength properties are not constant across time.
  • The number of authors with nonzero instrength is plotted alongside total authors as functions of references and average publication year.The time coordinate represents the average publication year of papers in each dynamical slice.
  • The total number of network authors increases over time, while the number of cited authors grows more slowly.This pattern reflects the increasing production of scientists in physics.
  • Rescaled indegree distributions from different networks obey the same universal curve.The rescaled variable is k_in/⟨k_in⟩, where k_in is the number of citing authors and ⟨k_in⟩ is its network-wide average.
  • Instrength distributions do not show universal behavior after a simple scale transformation.

IV. SCIENCE AUTHOR RANK ALGORITHM

SARA ranks authors by iteratively diffusing scientific credit through a weighted author network. The algorithm combines link-weighted credit transfer with redistribution and productivity-normalized credit allocation across coauthors.

  • SARA uses global network structure to rank authors by iteratively diffusing scientific credits.Each author begins with one unit of credit, distributes it to neighbors according to directed-link weights, and receives credit for redistribution in later iterations.
  • Credit transferred from author j to author i is proportional to the directed edge weight w_ji.The score equation also includes a damping factor q and redistribution terms.
  • A portion q of each node’s credit is redistributed to all scientists rather than only following network links.
  • The productivity factor z_i is computed from normalized credit assigned to an author across papers and coauthor counts.Each paper’s credit is redistributed equally among its coauthors, so papers carry equal knowledge weight rather than authors receiving homogeneous baseline credit.

A. Ranking Authors

The authors use SARA to track authors’ relative ranks through time and compare them with established citation-based rankings. Nobel laureates reach top relative performances near their award dates, while dynamical slices account for changing citation rates.

  • Ranking Authors: SARA scores determine author rank, and dynamical slices with equal citation counts track rank evolution across historical periods.
  • Ranking Authors: Relative rank measures the proportion of authors with better scores and therefore expresses each scientist’s top percentile.
  • Ranking Authors: SARA and citation-based rankings can differ substantially, with prominent physicists sometimes ranking far higher in SARA than in CC or BCC.The figure identifies differences using scatter plots across 1893–1966 and 2005 networks.
  • Ranking Authors: Relative rank is preferred for historical comparisons because the number of authors in the network increases rapidly over time.
  • Ranking Authors: Nobel laureates’ relative-rank dynamics are qualitatively related to their prize dates, with top performances reached close to the awards.
  • Ranking Authors: Equal-citation dynamical slices account for increasing citation rates, making rank change with research activity and the success of new research fronts.

V. COMPARISON WITH DIFFERENT METRICS

The paper compares SARA with Citation Count and Balanced Citation Count, two author-ranking measures that differ in how they treat citation information.

  • Compared metrics: Citation Count ranks authors by total citations received within a time window.The paper notes that the citation count is not identical to author indegree in the citation network.
  • Compared metrics: Balanced Citation Count normalizes citation weight by the number of authors on the cited paper.Authors are ranked using their instrength under this normalization.
  • Comparison design: SARA is evaluated alongside these two commonly used ranking schemes to assess author-impact rankings.The comparison motivates the subsequent test of predictive power using major physics prizes.

VI. BENCHMARKING THE SCIENCE AUTHOR RANK ALGORITHM

The paper benchmarks SARA against major physics awards and compares its ranked scientists across historical and recent publication periods. SARA places award winners more consistently in top rank brackets, while its historical top-20 list overlaps substantially with later prize recipients.

  • Award benchmark: 77% of considered honors went to scientists whose best SARA performance rank was below 1%, compared with 66% for Citation Count and 67% for Balanced Citation Count.SARA also placed 35% of prizes below the 0.1% rank threshold.
  • Award benchmark: The benchmark measures each prize winner by their best historical performance percentile, where a lower percentile indicates better performance.The percentile is the percentage of authors with a better rank at the same time.
  • Award benchmark: SARA predicts major physics prize assignments better than Citation Count and Balanced Citation Count.The benchmark treats awards as an outcome of peer evaluation and examines whether award winners occupy top rank brackets.
  • Top-ranked scientists: 16 of 20 scientists listed for 1967–1973 had earned one of the considered major prizes.The table separately lists top-20 SARA scientists for 1967–1973 and 2003–2004 who had not yet received those awards.
  • Top-ranked scientists: All scientists listed for 2004 were described as top physicists in their fields and probably eligible for important physics prizes under the paper’s criteria.The authors caution that some prizes are disciplinary and therefore do not apply to every author.

VII. CONCLUSIONS

The paper presents SARA as a diffusion-based scientist-ranking measure that captures non-local scientific credit spreading. It also emphasizes that the method’s conclusions are constrained by its dataset and treatment of citations as positive impact indicators.

  • Conclusions: SARA ranks scientists by mimicking the spread of scientific credits among authors.It combines a biased random walk with random credit distribution across network nodes.
  • Conclusions: The algorithm gives greater importance to inlinks from highly ranked authors and captures non-local credit-spreading effects.Any author can in principle receive scientific credit through the network’s diffusion process.
  • Limitations: The methodology includes caveats associated with adopting ranking approaches without critical examination.The conclusion introduces these caveats before discussing dataset scope and citation interpretation.
  • Limitations: The authors caution that SARA uses only the Physical Review dataset, which can undervalue scientists with substantial impact outside physics.They suggest broader repositories could mitigate this limitation, while increasing repositories would intensify author-disambiguation problems.
  • Limitations: The credit-spreading model treats credits and citations as positive indicators of impact, although the interpretation of negative citations is debated.This limitation concerns how citation valence is represented in the ranking.

Appendix A: IDENTIFICATION AND DISAMBIGUATION OF AUTHORS

The appendix describes author identification in the WACN and examines how SARA rankings vary with the damping factor q. It also reports identifier ambiguity and prize-based evidence supporting q = 0.1.

  • Author identification: The WACN construction uses author identifiers formed from the full last name and the initials of first and middle names.The paper illustrates this convention with Einstein and Bethe examples.
  • Author identification: 216 623 different authors are distinguished using this identification approach.
  • Author disambiguation: Author identification is limited because name variations can prevent matching one person, while identical initials and surnames can merge distinct people.The paper does not apply more elaborate disambiguation.
  • Damping-factor dependence: SARA rankings for different q values are linearly correlated, with weaker correlation as the difference between q values increases.The comparisons use rankings from the WACN built from papers published between 1893 and 1966.
  • Damping-factor dependence: q = 0.1 is selected because most prize winners reached top SARA positions, and SARA was more effective than CC or BCC for predicting future prize winners.The prize intervals used in Figure 10 are arbitrary, although the reported results do not strictly depend on that choice.
Loading 0907.1050v2…