Source-linked AI summary

Caveats for the journal and field normalizations in the CWTS ("Leiden") evaluations of research performance

Tobias Opthof, Loet Leydesdorff

arXiv:1002.2769v1cs.DLphysics.soc-ph

TL;DR

The paper examines problems arising from citation-score normalizations based on averaging before division, especially for skewed distributions. It contrasts this with dividing publication-level observed and expected values first, reports a corrected score of 0.91 (± 0.11), and argues for greater transparency in using these indicators.

  • Problem

    Skewed citation-score distributions can make normalization based on averaging first and dividing thereafter produce large effects.

  • Method

    The paper compares averaging first and dividing thereafter with first dividing observed and expected values and then averaging.

  • Results

    0.91 (± 0.11) is the score obtained when the proposed normalization is performed.

  • Takeaways & Limitations

    The paper encourages evaluators to specify their normalization decisions and improve the transparency and traceability of indicators increasingly used in institutional management.

  • Takeaways & Limitations

    Confidentiality prevents critical users of CWTS evaluation reports from correcting the normalization, although disaggregated tables would allow such corrections.

Abstract

from arXiv · show

The Center for Science and Technology Studies at Leiden University advocates the use of specific normalizations for assessing research performance with reference to a world average. The Journal Citation Score (JCS) and Field Citation Score (FCS) are averaged for the research group or individual researcher under study, and then these values are used as denominators of the (mean) Citations per publication (CPP). Thus, this normalization is based on dividing two averages. This procedure only generates a legitimate indicator in the case of underlying normal distributions. Given the skewed distributions under study, one should average the observed versus expected values which are to be divided first for each publication. We show the effects of the Leiden normalization for a recent evaluation where we happened to have access to the underlying data.

A real-life example

A real-life reconstruction for one biomedical-engineering principal investigator shows that normalization choices can materially change performance scores and managerial interpretation. The proposed publication-level normalization yields a score near the world average, unlike the CWTS values, but document-type and measurement choices constrain the comparison.

  • Methodological issue: Citation distributions are far from normal, so averaging before division can differ substantially from dividing publication-level observed and expected values first.The operation is therefore sensitive to the order of averaging and division; effects may be smaller for more nearly normal citation distributions.
  • Study design: The example covers 65 articles by a leading Amsterdam Medical Center scientist in biomedical engineering during 1997–2006.The analysis also noted one review article and two editorials, while the biomedical-engineering subfield includes engineering and medical applications.
  • Study design: The authors used articles and proceedings papers together, did not correct for self-citations, and limited analysis to JCS because FCS could not be computed.They report that normalizing with articles only produced almost identical results, and the CWTS report did not provide JCSm+ including self-citations.
  • Results: 0.91 (± 0.11) was the proposed normalization score, not significantly different from the world average.The authors state that this calculation is provided in the Appendix.
  • Results: 0.71 with self-citations and 0.58 without them were the CWTS values, indicating significantly below-world-average performance for the researcher.The proposed score would have given institutional management no reason for concern, whereas the CWTS value raised questions.
  • Implications: A 0.20 absolute difference equals the gap between the world average of 1.00 and the CWTS underperformance threshold of 0.80.The authors question using an indicator that varies by more than 28% for a single scientist in managerial or policy purposes.

Ranking the evaluated scientists

The study asks whether replacing the CWTS normalization changes rankings among 232 scholars and whether rank differences are statistically meaningful. It focuses on how citation-distribution shapes and sample sizes affect those comparisons.

  • Research questions: The authors ask whether equal CWTS ratings produce different values under their model and how changes affect high- and low-ranked scientists.They also ask whether rank differences are significant given the underlying distributions.
  • Assumption: The effect is expected to be smaller when citation distributions across documents are more nearly normal.This is stated as an assumption about the relationship between distribution shape and normalization effects.
  • Data and approach: The dataset’s citation counts were N = 223,425; x = 963 ± 1,138; median = 566.The authors lacked automated access to journal and field normalizations and therefore studied the ranking question through central ranking issues.

CPP/JCSm (CWTS)

The CWTS and proposed normalizations often produce similar overall rankings but can diverge substantially for individual scientists. Statistical testing shows that significance depends on the underlying citation distributions and sample sizes, not on visual ranking differences alone.

  • Overall comparison: The two methods were highly correlated, with r = 0.99 and 0.94, respectively; p < 0.01.Despite this correlation, the authors report important quantitative differences in some cases.
  • Individual ranking cases: 117th and 118th scientists tied at CPP/JCSm = 1.00 in CWTS data remained tied at 1.03 in replication but scored 0.93 and 1.50 under the authors’ method.This illustrates that identical or near-identical CWTS values can separate under publication-level normalization.
  • Statistical testing: Only the second and third authors differed significantly from 1, whereas the first did not.The authors tested distributions against the world-average value of 1 using their ratio-based results.
  • Statistical testing: Some pairwise differences, including between ranking extremes, were significant after Bonferroni correction, while others were not.The significance of differences depends on distribution shape and sample size and cannot be inferred from tables or graphs alone.
  • Individual ranking cases: Two authors tied at 1.00 in Leiden data obtained very different values under the authors’ indicator, yet their difference was not significant.One of these authors was not significantly different from the top author despite Leiden ratings of 1.00 and 2.18.
  • Systematic pattern: The normalization difference was not a fixed percentage and tended to become larger as the CWTS indicator decreased.The authors interpret this pattern as systematic underrating of low-ranked scientists in the CWTS evaluation.

Conclusions and discussion

The authors argue that Leiden indicators have consequential effects on institutional management, especially as they increasingly inform faculty, departmental, promotion, and funding decisions. Because users often receive only above/below-average management information and cannot correct the normalization, the authors call for transparent, traceable, and explicitly justified use of these measurements.

  • Leiden indicators increasingly inform micro-management, promotion, and funding decisions at faculties and departments.
  • The availability of these indicator constructs affects institutional management.
  • Critical users of CWTS evaluation reports cannot correct for the problems associated with the normalization because the underlying information is confidential.
  • Users obtain only management information indicating whether a research group is above or below average under the applied normalization and summary statistics.
  • The authors encourage scientometric evaluators to specify their decisions in applying measurements and normalizations and to ensure indicator transparency and traceability.
Loading 1002.2769v1…