Source-linked AI summary
The use of percentiles and percentile rank classes in the analysis of bibliometric data: Opportunities and limits
Lutz Bornmann, Loet Leydesdorff, Ruediger Mutz
TL;DR
The paper examines percentiles as alternatives to mean-based citation indicators and addresses unresolved choices in calculating percentiles and percentile rank classes. It describes their use in bibliometric evaluation and concludes that percentile rank classes are useful, while methodological consistency remains necessary for comparing studies.
Problem
Mean-based citation indicators are limited for skewed data, and percentile calculations and rank assignments lack a single satisfactory solution.
Method
The study examines opportunities, limits, and calculation problems associated with percentiles and percentile rank classes in bibliometrics.
Results
Percentile rank classes are very useful for evaluating research performance.
Takeaways & Limitations
Results from percentile-based studies are difficult to compare unless they use the same choices for percentile calculation and rank assignment.
Takeaways & Limitations
The study’s reviewed approaches to percentile and percentile-rank-class calculation are not yet completely satisfactory.
Abstract
from arXiv · showhide
Percentiles have been established in bibliometrics as an important alternative to mean-based indicators for obtaining a normalized citation impact of publications. Percentiles have a number of advantages over standard bibliometric indicators used frequently: for example, their calculation is not based on the arithmetic mean which should not be used for skewed bibliometric data. This study describes the opportunities and limits and the advantages and disadvantages of using percentiles in bibliometrics. We also address problems in the calculation of percentiles and percentile rank classes for which there is not (yet) a satisfactory solution. It will be hard to compare the results of different percentile-based studies with each other unless it is clear that the studies were done with the same choices for percentile calculation and rank assignment.
1 Using reference sets
Reference sets normalize citation impact by comparing publications with similar subject area, year, and document type, but mean-based indicators are problematic for skewed citation distributions. Percentile-based approaches offer an alternative because they reduce sensitivity to extreme values and support comparisons across heterogeneous publication sets.
- Limits of reference sets: Other citation-related factors, including author count and publication length, can influence citations systematically but are not included when reference sets are created.
- Reference-set normalization: Reference sets group publications by subject area, publication year, and document type to estimate expected citation impact.Mean observed citation rate divided by mean expected citation rate yields the Relative Citation Rate (RCR).
- Reference-set normalization: Normalized indicators make citation impact comparable across fields, publication years, and document types.
- Limits of means: Citation distributions are usually skewed, so a few highly cited publications can determine the reference-set arithmetic mean.The arithmetic mean is therefore unsuitable as a measure of central tendency for such data.
- Limits of means: The Göttingen ranking illustrates how one extremely highly cited publication can strongly influence a mean-based quotient.The paper reports that this influence supported an excellent or outstanding attribution based on erroneous assumptions about the citation distribution.
2 The calculation of percentiles
Percentiles and percentile rank classes provide alternatives to mean-based citation indicators, especially for skewed data and extreme values. Their calculation requires explicit choices about ranking, scale boundaries, and percentile formulas, and these choices affect comparability across analyses.
- Purpose and advantages: A percentile is a value below which a certain proportion of observations fall, enabling classification among the most-cited publications.
- Purpose and advantages: Percentiles are less influenced by extreme values than means and do not require choosing a probability density function.
- Applications: Percentile rank classes support classification and comparison of publication shares across different publication sets.
- Calculation procedure: Percentile calculation ranks publications by decreasing citations, assigns average ranks to ties, sets scale boundaries, and applies a chosen formula.For example, two publications with 15 citations receive rank 40.5 rather than ranks 41 and 40.
- Calculation procedure: Different formulas can produce different percentile values and may fail to give the median a percentile of 50 or treat distribution tails symmetrically.For 41 publications, the formula 21/41*100 gives 51.22, while ((i-1)/n*100) gives a mean percentile of 48.78 in the example.
- Applications: Percentile distributions can compare institutional citation impact, with violin plots showing that Univ 3 has lower median inverted percentiles and fewer zero-citation publications than the others.
3 Assigning percentiles to percentile rank classes
Percentiles can be assigned to percentile rank classes to interpret and compare citation impact across publications and publication sets. Several class schemes are used, including schemes emphasizing highly cited papers and top-citation distinctions.
- Percentile rank classes make publication-level percentiles interpretable for evaluating citation impact.
- PR(2,10) separates publications below the 90th percentile from those at or above it, identifying the 10% most cited papers.
- PR(6) groups percentiles into six classes, from papers below the 50th percentile through the top 1%.
- The ESI scheme uses six thresholds from the top 50% to the top 0.01% and provides finer distinctions among highly cited publications.
- PR distributions compare publication sets with expected shares from random database selections, revealing differences across universities.
- Univ 3 has fewer papers below the 50th percentile and more papers in the 5% and 1% classes than the other universities.
4 Rules for and problems with assigning publications to percentile
Creating percentile rank classes requires rules about class numbers and reference-set size, but tied citation counts can prevent unambiguous assignments. Proposed remedies improve aggregate proportions while introducing analytical trade-offs.
- The number of classes should not exceed the number of publications, and PR(6) requires at least 101 untied papers in one reference set.
- Journal-based reference sets may be too small to create some percentile classes, especially for journals publishing many reviews.
- Tied citation counts can make class assignment ambiguous; in the example, ties would place 15% rather than 10% of papers in the 10% class.
- Citation disparity, which affects assignment difficulty, varies with subject category, citation-window length, and reference-set size.
- Fractional assignment ensures exactly 10% and 90% aggregate shares for one reference set, but individual papers no longer have unambiguous classes.
- Fractional assignments permit more precise class proportions but cannot support paper-level analyses that require unambiguous outcomes, such as logistic regression.
5 Discussion
Percentiles and percentile ranks offer useful alternatives to mean-based bibliometric indicators, but their calculation and class assignment involve unresolved choices that constrain comparability across studies.
- Opportunities: Percentiles and percentile ranks are established alternatives to mean-based indicators for normalized citation-impact evaluation.The study presents percentile ranks as useful in evaluation studies, including comparisons between research institutions.
- Limits: Percentiles can be calculated in different ways, including the frequently used Hazen formula and the possible Gringorten approach.Further studies are needed to assess the advantages and disadvantages of these approaches for bibliometric data.
- Opportunities: Percentile ranks can reveal differences in performance between research institutions such as universities.
- Limits: Reference-set size constrains class design: the number of classes should not exceed the number of papers plus one.
- Limits: Tied citation counts create ambiguous assignments of publications to percentile-rank classes, and existing remedies remain unsatisfactory.The study seeks procedures that assign publications to classes using unambiguous citation thresholds where possible.
- Limits: Results from percentile-based studies are difficult to compare unless they use the same percentile-calculation and rank-assignment choices.The paper therefore emphasizes caution when applying percentile techniques.