Source-linked AI summary

The h-index is no longer an effective correlate of scientific reputation

Vladlen Koltun, David Hafner

arXiv:2102.03234v1cs.DL

TL;DR

The paper examines whether citation-based measures remain reliable indicators of scientific reputation as publication patterns change. Using large datasets across four fields, two bibliographic platforms, and award-based recognition benchmarks, it finds declining h-index effectiveness and identifies fractional measures—especially h-frac—as more robust alternatives.

  • Problem

    The study asks whether citation-based measures, especially the h-index, remain reliable indicators of scientific reputation as publication patterns and authorship change.

  • Method

    The authors analyze millions of articles and citations across four fields and two bibliographic platforms, correlating metric-based rankings with rankings based on scientific awards.

  • Results

    The h-index’s award correlation in physics fell from 0.34 in 2010 to 0.00 in 2019, while fractional measures improved effectiveness and h-frac consistently outperformed alternatives.

  • Takeaways & Limitations

    The findings suggest reconsidering h-index-based rankings and using fractional measures such as h-frac as more robust alternatives.

  • Takeaways & Limitations

    The study does not address cross-field normalization, self-citations, or whether author order should determine credit allocation.

Abstract

from arXiv · show

The impact of individual scientists is commonly quantified using citation-based measures. The most common such measure is the h-index. A scientist's h-index affects hiring, promotion, and funding decisions, and thus shapes the progress of science. Here we report a large-scale study of scientometric measures, analyzing millions of articles and hundreds of millions of citations across four scientific fields and two data platforms. We find that the correlation of the h-index with awards that indicate recognition by the scientific community has substantially declined. These trends are associated with changing authorship patterns. We show that these declines can be mitigated by fractional allocation of citations among authors, which has been discussed in the literature but not implemented at scale. We find that a fractional analogue of the h-index outperforms other measures as a correlate and predictor of scientific awards. Our results suggest that the use of the h-index in ranking scientists should be reconsidered, and that fractional allocation measures such as h-frac provide more robust alternatives. An interactive visualization of our work can be found at https://h-frac.org

Introduction

The study evaluates scientometric measures against scientific awards using large, cross-field datasets and finds declining h-index effectiveness, with fractional allocation improving robustness and h-frac outperforming alternatives.

  • The study analyzes 1.3 million articles and 102 million citations across biology, computer science, economics, and physics using Scopus and Google Scholar.It evaluates the 1,000 most highly cited researchers in each field.
  • Scientific awards serve as recognition-based benchmarks for comparing rankings induced by citation measures across four fields and two bibliographic platforms.The awards include Nobel Prizes, Breakthrough Prizes, national academy membership, and field-specific distinctions.
  • The study extends prior work through more than 10,000 awards, 1,848 distinct awards, 4,000 scientists, and yearly data from 1970 onward.This temporal granularity enables evaluation of how metric effectiveness and predictive power change over time.
  • The h-index’s correlation with scientific awards in physics declined from 0.34 in 2010 to 0.00 in 2019, alongside changing authorship patterns and increased hyperauthorship.The reported values use Kendall’s τ in the Scopus physics dataset.
  • Fractional citation allocation mitigates declining metric effectiveness, and h-frac consistently outperforms alternative measures as a correlate and predictor of awards.Controlled experiments across datasets support the robustness of these findings.

Results

Traditional scientometric measures, especially the h-index, have become less effective as indicators of scientific recognition amid rising coauthorship. Fractional citation allocation mitigates this decline, with h-frac performing best across datasets and robustness conditions.

  • Declining effectiveness of individual research metrics: By 2019, 68% of highly cited physicists averaged more than 100 coauthors per publication, and hyper-collaborators had entered the h-index ranking.In 1980, 84% averaged fewer than 10 coauthors per publication.
  • Fractional allocation: Across fields and data platforms, h-frac was the strongest correlate of scientific awards, with average τ = 0.32 in 2019 versus 0.16 for the h-index.Measures normalized by coauthor count generally outperformed unnormalized indicators.
  • Predictive power and other measures: h-frac had the highest five-year predictive power across datasets, remaining stable at average τ values of 0.34, 0.36, and 0.33 in 1994, 2004, and 2014.The h-index’s predictive power declined from average τ = 0.32 in 2004 to 0.24 in 2014.
  • Robustness of the findings: Fractional measures consistently outperformed traditional counterparts, and h-frac remained the most reliable indicator across alternative statistics, awards, researchers, and counting conditions.This robustness held even when restricting researchers to those averaging at most 100 coauthors per paper.
  • Further analysis: Hyperauthors commonly exceeded h-index values of 80 by 2019, whereas their h-frac values remained predominantly below 20.High h-frac rankings still included prolific collaborators, including scientists averaging 4.3–5.6 authors per publication.

Discussion

The study finds that conventional scientometric measures have become less effective indicators of scientific reputation, while fractional allocation—especially h-frac—improves robustness across conditions.

  • h-frac is the most reliable measure across different experimental conditions.
  • Fractional allocation neutralizes hyperauthorship’s inflationary effects while continuing to reward impactful collaborative research.
  • The study does not address cross-field normalization, self-citations, or author order in credit allocation.
  • The authors recommend reconsidering the h-index and using h-frac as a more robust alternative for assessing individual scientific impact.
  • An interactive visualization of the study is available at h-frac.org.

Methods

The study constructs four-field researcher datasets by matching Google Scholar and Scopus profiles, cleaning publication records, and collecting citation, authorship, and award information.

  • Researchers are retrieved from Google Scholar, matched to Scopus profiles, filtered by primary subject area, and reduced to the top 1,000 per field.
  • For all 4,000 researchers, the study collects publication year, author count, and yearly citation data from Google Scholar.
  • Google Scholar under-reports authors in large teams, but consistency with Scopus indicates that findings are robust to this limitation.
  • Scopus records are collected for the same researchers and require no special cleaning because they are significantly less noisy.
  • Google Scholar contains approximately twice as many publications and citations as Scopus, reflecting broader but noisier coverage.
  • The study identifies awards as reputation indicators, manually verifies candidate laureate matches, and retains award years for temporal analyses.

Supplementary Information

The supplementary datasets contain millions of publications and citations from Google Scholar and Scopus, alongside award records for the 4,000 researchers.

  • Google Scholar contains 2,624,994 valid publications cited 220,783,854 times, with yearly coverage from 1970 onward.
  • Scopus contains 1,290,219 publications and 102,405,086 citations across biology, computer science, economics, and physics.
  • The award dataset traces 1,848 distinct awards to the 4,000 researchers, with 976 researchers receiving at least one award.
  • The supplementary material introduces the scientometric measures evaluated in the study, including the h-index and citation-based alternatives.

µ-Index

The study defines several traditional citation measures and constructs fractional counterparts by normalizing citation counts by the number of authors.

  • The g-index is the largest g for which g publications collectively have at least g^2 citations, while the o-index combines h with the most-cited publication’s citations.
  • The m-index is the median citation count among the publications in the h-core, defined as the top h publications by citation count.
  • Fractional measures normalize each publication’s citation count by its number of authors, distributing credit equally without introducing new parameters.
  • The fractional h-index, h-frac, is defined from normalized citation counts ranked in decreasing order.
  • The c-index aggregates all citations, and c-frac aggregates citation counts after author-based normalization.

c-frac =

These fractional measures are defined by applying normalized citation counts to established bibliometric constructions. The supplied passages specify μ-frac as a mean and describe analogous fractional variants for the g-index and o-index.

  • μ-frac is the mean of normalized citation counts across all publications by an author.
  • g-frac is defined by analogy with the g-index using normalized citation counts.
  • o-frac is defined as the geometric mean of h-frac and the largest normalized citation count.

o-frac =

The evaluation ranks scientists by each scientometric measure and compares those rankings with award-based recognition. ROC curves and AUC quantify how effectively measures prioritize award recipients, with h-frac performing best across the studied datasets.

  • The ROC analysis ranks scientists by each measure, then aggregates awards as scientists are added in rank order.
  • A higher AUC indicates that a measure ranks more award-winning scientists near the top.
  • h-frac performs best across all research areas and datasets in the ROC analysis.
  • The evaluation also compares measure-based rankings with award rankings using correlation criteria.

Kendall’s τ

The study evaluates ranking agreement with award-based rankings using Kendall’s τ and Somers’ D, while accounting for ties and directional asymmetry. These statistics characterize concordance between scientometric measures and scientific recognition.

  • Kendall’s τb accounts for ties when comparing two rankings.
  • Kendall’s τ uses concordant and discordant pairs together with ties occurring in either ranking.
  • When no ties are present, Kendall’s τb reduces to τa.
  • Somers’ D is asymmetric, with the scientometric ranking treated as A and the award ranking as B.

Spearman’s ρ

The study also uses Spearman’s ρ to compare scientometric and award-based rankings. Across the reported correlation statistics, fractional measures are consistently stronger, with h-frac overall the most effective measure, while economics shows smaller differences and favors g-frac and o-frac.

  • Spearman’s rank correlation is the Pearson correlation coefficient between the two rank variables.
  • h-frac is overall the most effective measure for correlation with award-based scientific reputation.
  • Fractional measures consistently outperform their non-fractional counterparts across correlation statistics.
  • In economics, g-frac and o-frac appear most effective, but differences among measures are smaller than in other fields.

Temporal Dynamics

Scientometric measures generally lose effectiveness and predictive power over time, while fractional measures remain more stable and increasingly outperform traditional measures. h-frac is the most effective correlate and predictor of award-based scientific reputation.

  • Temporal Dynamics: The study evaluates effectiveness over time using publication, citation, and award data available through each year.This design supports temporal analysis of both correlation and predictive power.
  • Temporal Dynamics: From 2014 onward, all fractional measures are more effective than traditional measures, with h-frac the strongest correlate of scientific awards.The gap between fractional and non-fractional measures increases over time.
  • Temporal Dynamics: Most scientometric measures decline in effectiveness over time, whereas fractional measures are more stable and their advantage grows.Effectiveness is measured by correlation between metric-based rankings and rankings induced by scientific awards.
  • Temporal Dynamics: From 2014 onward, all fractional measures are more predictive than traditional measures, and h-frac is the most predictive measure.The analysis evaluates prediction of awards five years into the future.

Additional References

The cited references define standard rank-correlation measures used to evaluate scientometric rankings and association.

  • Additional References: Kendall’s references cover tie treatment and rank correlation, while Somers’ D, Goodman–Kruskal’s γ, and Spearman’s rank correlation provide related association measures.These references correspond to evaluation criteria used in the study.
Loading 2102.03234v1…