Source-linked AI summary

Turning the tables in citation analysis one more time: Principles for comparing sets of documents

Loet Leydesdorff, Lutz Bornmann, Rüdiger Mutz, Tobias Opthof

arXiv:1101.3863v2cs.DLphysics.soc-ph

TL;DR

The paper addresses limitations of averages-based citation indicators for skewed citation distributions and develops a percentile-rank framework. It organizes indicators by reference sets, evaluation criteria, and publication-set independence, yielding proposed measures that account for citation-distribution size and shape.

  • Problem

    Citation indicators commonly use arithmetic averages despite highly skewed citation distributions, while normalization choices and productivity effects remain insufficiently addressed.

  • Method

    The paper develops percentile-rank indicators organized around reference-set selection, evaluation criteria, and whether publication sets are independent.

  • Results

    The proposed indicators [R(6), R(100), R(6,k), R(100,k)] improve on averages-based indicators by accounting for citation-distribution size and shape.

  • Takeaways & Limitations

    Percentile ranks provide a scheme for comparing document sets across alternative reference sets, evaluation criteria, and productivity structures.

  • Takeaways & Limitations

    The choice of external reference sets still requires analytical grounding.

Abstract

from arXiv · show

We submit newly developed citation impact indicators based not on arithmetic averages of citations but on percentile ranks. Citation distributions are-as a rule-highly skewed and should not be arithmetically averaged. With percentile ranks, the citation of each paper is rated in terms of its percentile in the citation distribution. The percentile ranks approach allows for the formulation of a more abstract indicator scheme that can be used to organize and/or schematize different impact indicators according to three degrees of freedom: the selection of the reference sets, the evaluation criteria, and the choice of whether or not to define the publication sets as independent. Bibliometric data of seven principal investigators (PIs) of the Academic Medical Center of the University of Amsterdam is used as an exemplary data set. We demonstrate that the proposed indicators [R(6), R(100), R(6,k), R(100,k)] are an improvement of averages-based indicators because one can account for the shape of the distributions of citations over papers.

Introduction

The paper argues that citation indicators should address skewed citation distributions, reference-set choices, evaluation criteria, and publication productivity rather than relying only on existing averages-based schemes.

  • Problem: Existing crown indicators rely on arithmetic averages even though citation distributions are typically highly skewed.This motivates considering percentile ranks and other non-parametric approaches.
  • Approach: The study develops percentile-rank indicators that normalize each paper against a reference set before aggregating results.The framework is designed to account for productivity when comparing citation distributions among document sets.
  • Problem: The paper identifies unresolved concerns about citation normalization, including potentially unsuitable journal classifications and heterogeneous journals.The cited concerns include errors in subject-category attributions and variation in document types, citation half-lives, and cognitive substance.
  • Problem: Existing indicators do not quantitatively assess trade-offs between papers at different high-citation thresholds, such as one top-1% paper versus five top-5% papers.Such comparisons matter for bibliometric evaluation in policy-making and institutional management.
  • Approach: The proposed framework separates reference-set selection, evaluation schemes, and whether publication sets are treated as independent.These degrees of freedom support comparisons involving fields, funding agencies, percentile thresholds, and productivity.
  • Approach: The resulting criteria aim to compare document sets while preserving distinctions in citation-distribution shape and evaluation purpose.The paper presents this as a more abstract scheme for organizing citation indicators.

4. The indicator should provide the user, among other things, with a relatively

The paper develops percentile-rank indicators that separate reference-set choice, evaluation criteria, and publication-set independence. Applied to seven scientists, these indicators make distribution shape and publication-rate effects visible when comparing citation impact.

  • Data and reference sets: The seven scientists’ document sets comprise 241 non-overlapping papers evaluated against same-journal, same-year, same-document-type reference sets.The source documents came from principal investigators at the Academic Medical Center of the University of Amsterdam.
  • Percentile-rank method: Each paper is assigned a percentile rank within its reference-set citation distribution, producing a 100-value rank structure for each set.Percentiles count the percentage of papers with fewer citations, exclude ties from that count, and are rounded to integer classes.
  • Publication-rate adjustment: R(100,k) removes uncontrolled publication-rate effects from means and medians, allowing papers across sets to be compared directly.Under this indicator, two papers at the 39th percentile can weigh as much as one paper at the 78th percentile elsewhere.
  • Empirical implications: For Scientists 1 and 2, percentile weighting shows that higher productivity can otherwise reduce the appreciation of papers in similar or higher percentile classes.In the [95th; 99th[ class, Scientist 1’s three papers contributed 0.65, while Scientist 2’s four papers contributed 0.54.
  • Indicator framework: The indicator framework distinguishes three analytical choices: reference sets, evaluation criteria, and whether publication sets are independent.These dimensions allow separate decisions in evaluation procedures.

More refined statistics

The study develops percentile-rank indicators and refines their statistical use by varying percentile classes, reference sets, and assumptions about sample independence. Applied to seven scientists, these methods reveal significant citation-impact differences and account for distribution shape and sample size.

  • Statistical results: Scientist 1 leaves the top group after 100-percentile testing, while correcting for sample size places Scientist 6 at the top of the ranking.Without correction, Scientist 1 benefits from having fewer papers; the paper contrasts 23 papers with Scientist 2’s 37 and Scientist 6’s 65.
  • Indicator framework: The proposed indicators R(6), R(100), R(6,k), and R(100,k) use percentile ranks rather than citation averages.They are designed to abstract from the shape of citation distributions using non-parametric statistics.
  • Indicator framework: The indicator scheme varies three analytically independent choices: reference sets, evaluation criteria, and whether publication sets are independent.Different choices produce different eventual scores, so the study treats them as explicit design dimensions.
  • Implications: The choice of percentile classes is normative, and external reference sets require analytical grounding.The study therefore identifies reference-set selection and class specification as boundaries requiring explicit justification.
  • Implications: The resulting percentile scores account for differences in citation-distribution size and shape and enable comparisons among sets of different sizes.The authors argue that percentile-rank indicators improve on averages-based indicators because reference-set choice is separated from the evaluation scheme.
Loading 1101.3863v2…