Source-linked AI summary
On the calculation of percentile-based bibliometric indicators
Ludo Waltman, Michael Schreiber
TL;DR
Percentile-based indicators must classify publications within discrete, tie-heavy citation distributions without biasing comparisons across fields. The paper introduces a fractional approach and formally shows that it avoids field-specific bias while targeting the desired proportions.
Problem
Discrete citation distributions and tied citation counts make percentile-based indicators miss the intended top-10% proportion and reduce comparison accuracy across fields.
Method
The paper introduces a fractional approach that assigns publications portions of percentile intervals to handle ties.
Results
The formal framework shows that the approach produces indicators without bias in favor of or against particular fields.
Takeaways & Limitations
For PPtop 10%, the approach ensures exactly 10% top-10% publications in each field.
Takeaways & Limitations
Earlier threshold-normalization approaches can make a group’s indicator change because of citation changes elsewhere in the field.
Abstract
from arXiv · showhide
A percentile-based bibliometric indicator is an indicator that values publications based on their position within the citation distribution of their field. The most straightforward percentile-based indicator is the proportion of frequently cited publications, for instance the proportion of publications that belong to the top 10% most frequently cited of their field. Recently, more complex percentile-based indicators were proposed. A difficulty in the calculation of percentile-based indicators is caused by the discrete nature of citation distributions combined with the presence of many publications with the same number of citations. We introduce an approach to calculating percentile-based indicators that deals with this difficulty in a more satisfactory way than earlier approaches suggested in the literature. We show in a formal mathematical framework that our approach leads to indicators that do not suffer from biases in favor of or against particular fields of science.
1. Introduction
Percentile-based indicators assess publications by their position in a field’s citation distribution, addressing limitations of citation averages and motivating a tie-aware calculation approach.
- Percentile-based indicators evaluate publications by their position within a field’s citation distribution rather than their raw citation counts.
- Citation distributions are highly skewed, so average citations can be strongly influenced by a few exceptionally cited publications.
- The PPtop x% indicator measures the proportion of a group’s publications belonging to the top x% most frequently cited in their field.
- Discrete citation counts and tied citation values make top-10% classification ambiguous and can distort comparisons between fields.
- The paper introduces a tie-aware approach intended to avoid field-specific bias and ensure exactly 10% top-10% publications in each field.
2. Some empirical context
Real-world Web of Science citation distributions show that publications exactly at the top-10% threshold are common enough to complicate percentile calculations, especially in lower-citation fields.
- Across seven Web of Science fields, 86.9%–90.0% of publications fell below the top-10% threshold, while 9.6%–10.0% exceeded it.
- Between 0.4% and 3.6% of publications were exactly at the top-10% threshold.
- Threshold ties were most prevalent in economics and mathematics, the fields identified as having relatively low citation density.
- Table 1 reports, for each field, the top-10% threshold and percentages below, exactly at, and above that threshold.
3. Overview of different approaches to calculating percentile-based indicators
Earlier percentile approaches handle tied publications by binary, threshold-based, or fractional rules, but they either miss the exact target proportion or create counterintuitive field-comparison effects.
- The reviewed approaches include percentile assignment, threshold rules, and fractional treatment of tied publications.
- Binary approaches either include tied publications and produce 14.29% top-10% publications or exclude them and produce 4.76%.
- Earlier approaches fail to produce exactly 10% top-10% publications, reducing the accuracy of inter-field and intertemporal comparisons.
- Random-order fractional counting assigns half of each 10-citation publication in the example to the top 10%, yielding 9.52% top-10% publications.
- Threshold normalization can make a group’s indicator change when publications elsewhere in the field cross the threshold.
- Replacing two 10-citation publications with 11-citation publications raises the example group’s normalized result from 7.00% to 15.00%.
4. Alternative approach to calculating percentile-based indicators
The alternative approach assigns tied publications fractional membership in the top 10% so the indicator reaches exactly 10% while limiting threshold effects to groups with publications at that threshold.
- 4. Alternative approach to calculating percentile-based indicators: The approach assigns each publication a fractional share of the citation distribution and divides the threshold tie accordingly.For 105 publications, each publication represents 0.952% of the field’s citation distribution; the 10-citation segment is split so 5.24% enters the top 10%.
- 4. Alternative approach to calculating percentile-based indicators: Exactly 10% of publications are counted as top 10% by splitting the tied citation segment fractionally between the top and bottom portions.In the example, 0.550 of each publication with 10 citations is counted, yielding (0.550 × 10 + 5) / 105 = 10%.
- 4. Alternative approach to calculating percentile-based indicators: A one-publication shift from 10 to 9 citations raises the second group’s top-10% proportion from 5.50% to 6.11%.The assigned fraction for the remaining 10-citation publications increases to 0.611.
- 4. Alternative approach to calculating percentile-based indicators: Replacing two 10-citation publications with 11-citation publications lowers the second group’s top-10% proportion from 5.50% to 4.38%.The fraction assigned to the remaining 10-citation publications falls to 0.438.
- 4. Alternative approach to calculating percentile-based indicators: Small citation changes at the threshold affect only research groups that have publications at that threshold.Groups without threshold publications are unaffected, and groups cannot benefit from citation increases to publications they do not own.
5. Formal mathematical framework
The paper formalizes percentile-based indicators by assigning publication scores from percentile intervals, fractionally handling citation ties, and averaging those scores. It proves that the resulting indicators are independent of a field’s citation distribution and therefore avoid field-specific bias.
- Generic indicator: A generic indicator partitions the percentile range into N intervals, assigns scores to those intervals, and averages publication scores across a research group.The framework requires boundaries 0 = p0 < p1 < ... < pN = 1 and associated scores s1 < s2 < ... < sN.
- Special cases: PPtop 10% is recovered as a two-interval special case, while R(6) uses six intervals with parameter values specified in Table 2.PPtop 10% sets N = 2, p1 = 0.9, s1 = 0, and s2 = 1; R(6) uses parameters listed in Table 2.
- Field notation: The framework represents a field using ci, the number of publications with i citations, and qi, the proportion with fewer than i citations.The interval length [qi, qi + 1] equals the proportion of publications with exactly i citations.
- Formal score assignment: Each publication receives a score from its percentile interval, and publications spanning multiple intervals receive a weighted average based on interval overlap.The overlap fractions determine how a citation-count group is assigned across relevant percentile intervals.
- Proof of field independence: The formal framework proves that the proposed indicators do not depend on a field’s citation distribution, avoiding bias for or against particular fields.The proof establishes this by showing that the final expression depends only on percentile boundaries and interval scores.