Source-linked AI summary

A principal component analysis of 39 scientific impact measures

Johan Bollen, Herbert Van de Sompel, Aric Hagberg, Ryan Chute

arXiv:0902.2183v2cs.DLcs.CY

TL;DR

Citation counts and the Journal Impact Factor do not fully capture scientific impact in an increasingly online research environment. The paper applies PCA to 39 citation- and usage-based impact rankings, finding that impact is multidimensional and that measures separate into distinct rapid/delayed and popularity/prestige dimensions. The results place the JIF at the periphery rather than the core of this construct, while suggesting usage-based measures may better represent consensus.

  • Problem

    Existing citation-based indicators, especially the JIF, may not adequately represent scientific impact as scholarly activity increasingly generates online usage data and alternative measures.

  • Method

    The study performs PCA on rankings from 39 impact measures calculated from citation data, usage logs, and Scopus data, using pairwise Spearman rank correlations.

  • Results

    Scientific impact is multidimensional: the first three PCA components explain 92% of variance, while the first two distinguish rapid from delayed and popularity from prestige.

  • Takeaways & Limitations

    No single indicator adequately measures scientific impact; the JIF expresses a particular peripheral aspect, whereas usage-based measures such as Usage Closeness may better capture consensus.

  • Takeaways & Limitations

    The analysis may omit plausible measures, and additional components or nonlinear dimensionality-reduction methods could reveal further distinctions.

Abstract

from arXiv · show

The impact of scientific publications has traditionally been expressed in terms of citation counts. However, scientific activity has moved online over the past decade. To better capture scientific impact in the digital era, a variety of new impact measures has been proposed on the basis of social network analysis and usage log data. Here we investigate how these new measures relate to each other, and how accurately and completely they express scientific impact. We performed a principal component analysis of the rankings produced by 39 existing and proposed measures of scholarly impact that were calculated on the basis of both citation and usage log data. Our results indicate that the notion of scientific impact is a multi-dimensional construct that can not be adequately measured by any single indicator, although some measures are more suitable than others. The commonly used citation Impact Factor is not positioned at the core of this construct, but at its periphery, and should thus be used with caution.

Introduction

Scientific impact has traditionally been measured through citations, but online scholarly activity has produced many alternative citation-, usage-, and network-based measures whose relative suitability is unclear. The study compares 39 such measures using PCA to identify dimensions and clusters of scholarly impact.

  • Motivation: Citation counts are commonly used to quantify how publications influence subsequent work.Administrators and policymakers often rely on citation data for scientific-impact assessment.
  • Motivation: The Journal Impact Factor averages a journal’s citations over a two-year publication period, extending its use from journals to articles, authors, and institutions.Its availability and intuitive definition helped make it dominant despite extensively discussed undesirable properties.
  • Problem: The JIF remains widely used even though experts regard it as far from perfect because accepted alternatives are lacking.
  • Related measures: Researchers have proposed modified citation statistics, H-index variants, and social-network measures such as citation PageRank to address JIF shortcomings.
  • Related measures: Usage-log measures became feasible as online publishers, aggregators, and libraries recorded interactions at scales exceeding the number of existing citations.Elsevier reported 1 billion full-text downloads in 2006 versus approximately 600 million Web of Science citations.
  • Study aim: The study applies PCA to rankings from 39 plausible scholarly-impact measures derived from citation and usage data to identify major dimensions and clusters.The measures include 19 calculated from 2007 JCR data, 16 from MESUR usage logs, and 4 from Scimago Scopus data.

Methods

The study combines existing and newly calculated scientific-impact measures from citation and usage datasets, then analyzes the rankings they produce.

  • Methods: The methodology retrieves or calculates 39 scientific-impact measures from citation and usage datasets and analyzes correlations among their resulting rankings.

Data preparation and collection

The analysis draws on citation data, large-scale usage logs, and Scopus-based journal rankings assembled from multiple scholarly platforms and institutions.

  • Data sources: The citation source was the 2007 Journal Citation Reports Science and Social Science Editions.
  • Data sources: The usage source was MESUR’s collection of 346,312,045 interactions recorded between March 1, 2006 and February 1, 2007.Records came from publisher, aggregator, and institutional library portals.
  • Data sources: Additional journal rankings came from the Scimago project and were based on Elsevier Scopus citation data.
  • Processing: The researchers detailed procedures for retrieving and calculating the 39 measures and identified each measure with a unique table number.Those identifiers were used in subsequent diagrams and tables.

Retrieving existing measures

Existing measures included JCR citation statistics, Scimago rankings, and measures documented through the study’s data-processing and PCA tables.

  • JCR measures: The 2007 JCR provided citation-based measures for approximately 7,500 selected journals.
  • JCR measures: The Immediacy Index measures same-year average citations, while the Journal Impact Factor measures a two-year average per-article citation rate.
  • JCR measures: Citation Half-life is defined as the median age of articles cited by a journal in 2006.
  • PCA reporting: Table 1 reports varimax-rotated loadings for the first two PCA components and average Spearman correlations with all other measures.The five lowest average correlations are marked with stars.
  • Scimago measures: Scimago measures included Journal Rank, Cites per Doc, H-Index, and Total cites, all based on Scopus citation data.Their definitions use PageRank, average citations, h-index thresholds, or three-year citation totals, respectively.
  • Processing: Scimago rankings were downloaded, loaded into a MySQL database, and added four existing measures to the dataset.

Calculating social network measures of scientific impact.

The study constructs citation and usage networks from journal relations, then applies social-network measures that vary by network type and connection treatment.

  • Citation networks: Citation networks encode journal-to-journal citation counts for specified origin and target publication periods.The resulting matrix records observed citations from one journal and date range to another.
  • Network construction: The citation network contained 897,608 connections among 7,388 journals and had a density of 1.6%.It was built from 2006-origin citations pointing to publications from 2004 and 2005.
  • Usage networks: Usage networks are derived from journal clickstreams, with transition probabilities estimating how often one journal follows another in user sessions.The transition probability is based on observed successive journal appearances in MESUR clickstreams.
  • Network construction: The usage analysis used 346,312,045 interactions recorded across scholarly portals between March 1, 2006 and February 1, 2007.The data came from publishers, aggregators, and academic library systems.
  • Network measures: Four social-network measure classes—degree, closeness, betweenness, and PageRank—were applied to both citation and usage matrices.These measures capture different aspects of a journal’s structural position in each network.
  • Network measures: Measures varied according to weighted versus unweighted, directed versus undirected, and citation versus usage network configurations.The design produced eight possible variations per measure class, although some permutations were excluded when conceptually inappropriate.

Hybrid Measures

The study also includes hybrid and probability-based measures that combine established impact indicators or normalize citation and usage activity.

  • Hybrid measures: The Y-Factor multiplies a journal’s Impact Factor by its PageRank.It is listed among measures that do not fit the four principal social-network classes.
  • Probability measures: Journal Cite Probability is calculated from citation numbers listed in the 2007 JCR.It represents one of the additional measures outside the main social-network classes.
  • Probability measures: Journal Use Probability is the normalized frequency with which a journal is used in the MESUR usage-log data.This measure is based on usage rather than citation counts.
  • Hybrid measures: The Usage Impact Factor uses the JIF definition but expresses a journal’s two-year average article usage.It was treated as an additional measure in the analysis.

Measures overview

The analysis covers 39 impact measures spanning citation statistics, usage statistics, social-network measures, and alternatives to the Journal Impact Factor.

  • Measure set: 39 measures were analyzed, including 32 citation- and usage-based measures, four Scimago journal measures, and three 2007 JCR measures.The 32 network-based measures comprise 16 derived from citation matrix C and 16 from usage matrix U.
  • Measure set: The selected measures were intended to represent major classes of statistics and social-network measures proposed as alternatives to the JIF.The overview groups them into four major classes, beginning with citation and usage statistics.

Analysis

The authors correlate journal rankings, remove two measures lacking significant relationships with the others, and use PCA, clustering, and a two-component map to examine similarity among measures.

  • Analysis: Spearman rank-order correlations produced a 39 × 39 matrix comparing every pair of journal rankings.Correlations were calculated over the intersections of the available journal sets.
  • Analysis: Two measures—Citation Half-Life and the Usage Impact Factor—were removed because they lacked significant correlations with any other measure.Both had N = 39 and p > 0.05 for their correlations with the remaining measures.
  • Principal component analysis: 66.1% of variance was represented by PC1 and 17% by PC2, so the first two components covered 83.4% of variance in measure correlations.Adding PC3 increased coverage to 92.6%.
  • Principal component analysis: The authors projected measures onto PC1 and PC2 and applied varimax rotation to create an interpretable two-dimensional similarity map.Citation-based and usage-based measures are distinguished visually in Figure 2, with the JIF separately marked.
  • Analysis: Hierarchical and k-means clustering were used to cross-validate the PCA-based identification of measures producing similar journal rankings.Both analyses operated on the measure-correlation matrix.

Results and discussion

The PCA and clustering analyses reveal that scientific-impact measures occupy multiple dimensions, chiefly separating rapid from delayed impact and popularity from prestige. Usage measures form a distinct, internally consistent group, while the JIF and related normalized citation measures sit at the periphery of the broader impact construct.

  • Limitations and future research: The analysis is limited by the possible omission of other plausible measures and by projecting correlations onto only the two highest-ranked components.The authors propose examining additional component combinations, including lower-valued components, in future work.
  • PCA dimensions: 92% of ranking-correlation variance is explained by the first three PCA components, while exceeding 95% requires four components.The first two components alone explain 83.4%.
  • PCA dimensions: PC1 separates usage from citation measures and is interpreted as distinguishing rapid from delayed views of scientific impact.Citation Immediacy and two Citation Betweenness Centrality measures approximate the usage-side position.
  • PCA dimensions: PC2 separates citation statistics expressing popularity from citation social-network measures expressing prestige.The four broad clusters are usage measures, per-document citation popularity, total citation rates and distributions, and citation social-network measures.

Data files

The PCA's supporting ranking data are available upon request, except data obtained under proprietary licenses.

  • Ranking data supporting the Principal Component Analysis are available upon request, except data obtained under proprietary licenses.
Loading 0902.2183v2…