Source-linked AI summary

Integrated Impact Indicators (I3) compared with Impact Factors (IFs): An alternative research design with policy implications

Loet Leydesdorff, Lutz Bornmann

arXiv:1103.5241v2cs.DLphysics.soc-ph

TL;DR

The paper addresses the mismatch between citation impact and central-tendency statistics, proposing an integrated indicator based on percentile-normalized citation distributions. It shows that JASIST has higher impact than MIS Quarterly despite a lower impact factor, and presents I3 as a decomposable measure applicable across document sets and citation windows. The paper also notes that zero-citation papers do not contribute to I3.

  • Problem

    Central-tendency statistics inadequately represent impact because citation distributions can be highly skewed and impact is treated as a sum rather than an average.

  • Method

    I3 integrates citation distributions after normalizing papers to percentile ranks against a reference set, allowing impact to be compared and decomposed across document sets.

  • Results

    JASIST has higher impact than MIS Quarterly despite its lower impact factor, and percentile-distribution sums capture this difference when means and medians do not.

  • Takeaways & Limitations

    I3 provides a general citation-impact measure that can be applied across samples of different sizes and decomposed across analytical units.

  • Takeaways & Limitations

    Papers with zero citations do not contribute to I3.

Abstract

from arXiv · show

In bibliometrics, the association of "impact" with central-tendency statistics is mistaken. Impacts add up, and citation curves should therefore be integrated instead of averaged. For example, the journals MIS Quarterly and JASIST differ by a factor of two in terms of their respective impact factors (IF), but the journal with the lower IF has the higher impact. Using percentile ranks (e.g., top-1%, top-10%, etc.), an integrated impact indicator (I3) can be based on integration of the citation curves, but after normalization of the citation curves to the same scale. The results across document sets can be compared as percentages of the total impact of a reference set. Total number of citations, however, should not be used instead because the shape of the citation curves is then not appreciated. I3 can be applied to any document set and any citation window. The results of the integration (summation) are fully decomposable in terms of journals or instititutional units such as nations, universities, etc., because percentile ranks are determined at the paper level. In this study, we first compare I3 with IFs for the journals in two ISI Subject Categories ("Information Science & Library Science" and "Multidisciplinary Sciences"). The LIS set is additionally decomposed in terms of nations. Policy implications of this possible paradigm shift in citation impact analysis are specified.

Introduction to the problem

The paper argues that impact should be treated as a sum rather than a central-tendency statistic, because citation distributions can be highly skewed and differently sized journals can have different impact profiles. It proposes integrating citation curves after percentile normalization, showing that JASIST has higher impact than MIS Quarterly despite a lower impact factor.

  • Conceptual problem: JASIST has the higher impact despite an IF-2009 of 2.300 versus MIS Quarterly’s 4.485.The lower IF reflects JASIST’s tail of more than 300 additional publications with lower citation counts.
  • Conceptual problem: The 66 most-highly cited JASIST publications obtained 380 citations more than the 66 MIS Quarterly publications.This comparison illustrates why the lower impact factor does not identify the journal with the greater aggregate impact.
  • Conceptual problem: Impact is not an average: central-tendency statistics cannot capture increases when document sets are combined, so citation curves should be integrated instead of averaged.The authors note that directly integrating raw citation curves would reduce the measure to total citations without qualifying citedness.
  • Proposed approach: Percentile normalization allows citation distributions from unequally sized document sets to be compared while preserving distribution shape.The normalized distributions can be integrated using percentile-rank frequencies, combining publication quantity with normalized citedness.
  • Implications: I3 distinguishes higher aggregate impact that means and medians fail to capture, and its sums can be decomposed across analytical units.The indicator can be applied across document sets and time periods when a reference set is specified.
  • Empirical illustration: Using 100PR and 6PR, MIS Quarterly contributed 2.61% and 2.34% of LIS impact, whereas JASIST contributed 9.73% and 8.63%.JASIST was therefore considered the journal with the highest impact among the 65 LIS journals.

Methods

The study compares I3 with impact factors using percentile-normalized citation data from two ISI subject categories, with analyses of journal distributions and decomposable contributions.

  • Aggregation: Percentile ranks support comparisons across differently sized document sets and can be summed or decomposed by journals, nations, and other institutional units.Each paper’s contribution can also be expressed as a percentage of the reference-set impact.
  • Indicator construction: I3 assigns each paper a percentile rank within reference sets defined by document type, publication year, and ISI Subject Category.Tied citation counts receive the highest applicable percentile under the study’s counting rule.
  • Data: The analysis uses 5,737 citable items from 65 LIS journals and 24,494 citable items from 48 Multidisciplinary Sciences journals.Citation data were downloaded from the WoS in February 2011.
  • Statistical analysis: The study compares means, sums, standard errors, and correlations for percentile-based indicators and impact factors using journal-level analyses.Pearson and Spearman correlations are used to compare the new indicators with IFs.
  • Distribution testing: Dunn’s multiple-comparison procedure tests citation-distribution differences among journals with family-wise correction for the number of pairwise comparisons.For 50 journals, the study uses 1,225 comparisons and a significance threshold of 0.000041.

I3 for the 65 journals of LIS

For the 65 LIS journals, I3 combines publication volume with percentile-normalized citation rates and captures citation-distribution shape rather than only mean impact. Its rankings differ from IF rankings, with JASIST leading on I3 while MIS Quarterly ranks seventh.

  • Correlations: The I3 and total-citation correlation is .963 at the journal level, but the document-level correlation between citations and percentiles is .639.These results distinguish the strong aggregate association from the weaker document-level relationship.
  • Indicator properties: I3 incorporates both the number of publications and their citation rates, while retaining information from the tails of citation distributions.Unlike the h index, I3 does not discard lower-impact papers as irrelevant.
  • Journal rankings: JASIST ranks first on I3 and I3(6PR), whereas MIS Quarterly ranks seventh on both measures in the LIS comparison.JASIST is ranked first on every reported measure except the impact factor.
  • Journal rankings: Scientometrics ranks above the Journal of the American Medical Informatics Association despite having fewer citations and a lower impact factor.Scientometrics has 1,336 citations and an IF of 2.167, compared with 1,784 citations and an IF of 3.974 for the latter journal.

Multidisciplinary Sciences

In Multidisciplinary Sciences, I3 produces journal rankings that differ substantially from IF rankings and reveals citation-distribution differences not captured by average-based measures. Size differences strongly affect correlations, but percentile-based analyses distinguish journal impact patterns.

  • Dataset composition: Six major journals account for 65.2% of the 24,494 citable publications in the Multidisciplinary Sciences set.PNAS contributes 7,058 publications, Nature 2,285, and Science 2,253.
  • Citation distributions: PNAS has 27,419 more citations than Nature but an IF of 9.432, less than one-third of Nature’s despite its differently shaped citation distribution.The large tail of moderately cited papers is identified as disadvantaging the larger journal.
  • Rankings: I3 and IF correlate significantly across the 48 journals, yet their journal rankings are very different.The ranking divergence is reported despite the significant association between the two measures.
  • Rankings: Current Science and Chinese Science Bulletin move from IF ranks 22 and 20 to I3 ranks 5 and 6.Their IFs are 0.782 and 0.898, respectively.
  • Correlations: The partial correlation between I3 and publication count controlling for citations is .850, compared with -.724 for IF.For I3(6PR), the corresponding partial correlation is .982.
  • Citation distributions: Nature and Science differ significantly from all other journals and from each other, whereas PNAS is not significantly different from several journals.Forty-five of the 48 journals form a k=25 core set.

Performance measurement

I3 uses paper-level percentile ranks to decompose impact across journals, nations, institutions, and other units, enabling comparisons of productivity and impact distributions. In the LIS analysis, fractional country attribution and significance tests support these decompositions.

  • I3 assigns percentile ranks at the paper level, enabling impact decomposition across journals, nations, institutions, and other units.
  • Fractional counting assigns coauthored papers across countries while preserving total counts; one-third goes to country B and two-thirds to country A in the stated example.
  • The LIS country table includes units contributing at least 1% of total I3 and sorts them by I3 share relative to publication share.
  • The USA has the highest absolute contribution on both scales, while the Netherlands leads on I3 at 1.68 and Switzerland leads on I3(6PR) at 1.57.
  • I3(6PR) emphasizes highly represented top percentile segments, explaining Switzerland’s stronger position on that scale.
  • Percentile-based comparisons can test citation-curve differences with Dunn’s test and subset deviations from expectation with the Z-test.

Conclusions and discussion

The discussion argues that impact should be summed across paper-level percentile ranks rather than represented by central-tendency statistics. I3 incorporates citation-distribution size and shape, supports flexible reference sets and decompositions, and separates impact from statistical significance.

  • I3 requires a reference set and can be applied across document sets, time periods, databases, and citation sources.
  • Central-tendency statistics such as means and medians do not adequately capture highly skewed citation distributions.
  • I3 sums paper-level percentile ranks to integrate qualified citation curves, preserving both the size and shape of citation distributions.
  • The hundred-percentile scheme is general, can be represented continuously with decimal precision, and can be adapted to different policy contexts.
  • The six-percentile scheme enhances distinctions among highly cited units but can lose fine-grained distinctions when ranks are tied.
  • Impact differences should remain analytically separate from significance tests of citation distributions; non-significant differences are discouraged as policy inputs.
  • I3 is fully decomposable, allowing journals, nations, universities, institutions, authors, and other units to be evaluated within one framework.

Policy implications

The paper presents I3 as a policy-relevant alternative to IF-based evaluation because it handles skewed citation distributions and avoids disadvantaging productive units. Its rankings can differ materially from IF rankings, so statistical significance matters for policy use.

  • I3 avoids penalizing productive units whose additional less-cited papers depress average performance.
  • Mean percentile ranks swapped the first and sixth positions among seven evaluated principal investigators.
  • Policy decisions should avoid relying on differences that are not statistically significant.
  • Existing indicators often rely on parametric assumptions, whereas the paper generalizes a non-parametric approach for citation impact.
  • Current Science and Chinese Science Bulletin ranked fifth and sixth by I3 despite ranking 22nd and 20th by IF, respectively.
  • I3 is additive, allowing comparisons and aggregation of subsets that differ across journals, addresses, nations, or other dimensions.
  • Zero-citation papers do not contribute to I3, which combines publication quantity with normalized citation impact.
Loading 1103.5241v2…