Source-linked AI summary
Towards a new crown indicator: Some theoretical considerations
Ludo Waltman, Nees Jan van Eck, Thed N. van Leeuwen, Martijn S. Visser, Anthony F. J. van Raan
TL;DR
The paper examines whether the crown indicator’s field-normalization mechanism has a sound theoretical basis. Comparing it with an alternative mechanism, the authors find that the alternative has stronger theoretical properties and consistency, motivating a new crown indicator.
Problem
The paper addresses how citation counts should be normalized across fields in research-performance evaluations, where the crown indicator’s mechanism has been criticized.
Method
The authors theoretically compare the crown indicator’s and alternative normalization mechanisms using examples, consistency analysis, and treatment of overlapping fields.
Results
The alternative MNCS mechanism has a more solid theoretical basis and avoids counterintuitive results produced by the crown indicator in some examples.
Takeaways & Limitations
CWTS is moving toward MNCS as its new crown indicator, while recognizing that research performance should be assessed with complementary indicators.
Takeaways & Limitations
MNCS is generally not suitable for use in isolation and should be complemented with other indicators.
Abstract
from arXiv · showhide
The crown indicator is a well-known bibliometric indicator of research performance developed by our institute. The indicator aims to normalize citation counts for differences among fields. We critically examine the theoretical basis of the normalization mechanism applied in the crown indicator. We also make a comparison with an alternative normalization mechanism. The alternative mechanism turns out to have more satisfactory properties than the mechanism applied in the crown indicator. In particular, the alternative mechanism has a so-called consistency property. The mechanism applied in the crown indicator lacks this important property. As a consequence of our findings, we are currently moving towards a new crown indicator, which relies on the alternative normalization mechanism.
1. Introduction
The introduction motivates field-normalized citation indicators and outlines a theoretical comparison between the crown indicator’s normalization mechanism and an alternative based on averaging publication-level citation ratios. The paper examines their differences through fictitious examples, consistency, and the handling of overlapping fields.
- Motivation: Citation averages differ across scientific fields because of differences in referencing practices, cited-reference age, cross-field citation, and database coverage.These differences make raw citation counts difficult to compare across fields.
- Motivation: Careful field normalization is especially important for performance evaluations aggregated across countries, universities, or multidisciplinary research groups.CWTS uses a standard set of bibliometric indicators in performance evaluation studies.
- Normalization mechanisms: The crown indicator divides total actual citations by total expected citations, with expectations defined by document type, field, and publication year.Expected citations are based on the average citations of publications sharing those characteristics.
- Normalization mechanisms: The alternative mechanism averages each publication’s ratio of actual to expected citations rather than dividing aggregate actual citations by aggregate expected citations.This alternative was advocated by Lundberg (2007) and Opthof and Leydesdorff (2010).
- Paper approach: The paper theoretically compares the two mechanisms using fictitious examples, consistency analysis, and consideration of overlapping fields.The comparison is presented as the paper’s main analytical program.
2. Definitions of indicators
This section formally defines the CPP/FCSm and MNCS indicators, which differ in how they normalize citation counts. CPP/FCSm evaluates an oeuvre as a whole, whereas MNCS normalizes individual publications and gives each publication’s ratio equal weight.
- Indicator definitions: CPP/FCSm is CWTS’s established crown indicator, while MNCS is the new crown indicator CWTS plans to adopt.The indicators use different normalization mechanisms.
- CPP/FCSm rationale: CPP/FCSm treats a research group’s publications as one integrated oeuvre, so only total citations matter, not their distribution across publications.This rationale focuses on aggregate performance rather than independent publication-level performance.
- Normalization mechanisms: CPP/FCSm normalizes using a ratio of averages, whereas MNCS normalizes using an average of publication-level citation ratios.CPP/FCSm therefore operates at the oeuvre level, while MNCS operates at the individual-publication level.
- Indicator definitions: In exceptional cases where both citation and expected-citation counts are zero, the indicators define 0 / 0 = 1, treating the publication as having average performance.A nonzero citation count with a zero expected count cannot occur under the definitions.
- Normalization mechanisms: CPP/FCSm can be expressed as a weighted version of MNCS, assigning greater weight to ratios for publications with higher expected citation counts.Fields with higher average citations per publication consequently receive more weight than fields with lower averages.
3. Example 1
Example 1 shows that CPP/FCSm and MNCS can produce opposite overall rankings despite identical subfield-specific results. This difference arises because CPP/FCSm weights subfields unequally, whereas MNCS weights them equally, making MNCS invariant to certain field-specific rescaling.
- 3. Example 1: Research groups A and B have equal publication counts and identical shares in subfields X and Y, but they outperform one another in different subfields.The example asks which group has higher overall performance when group B leads in X and group A leads in Y.
- 3. Example 1: CPP/FCSm ranks research group A above B overall, even though B’s performance in each subfield exceeds A’s performance in the other subfield.The overall ranking is reported alongside the separate subfield values in Table 2.
- 3. Example 1: MNCS ranks research group B above A overall, while matching CPP/FCSm’s results within each subfield.Thus, the two indicators yield opposite overall rankings despite identical subfield-specific rankings.
- 3. Example 1: CPP/FCSm favors A because it gives more weight to subfield Y than X, whereas MNCS weighs both subfields equally.Both indicators treat a publication’s performance as the ratio of its actual to expected citations, but aggregate normalized performance differently.
- 3. Example 1: After citation normalization, the authors see no general reason to weight publications differently by field.They argue that indicators should correct field differences without subsequently treating publications from different fields differently.
- 3. Example 1: When actual and expected citations in subfield Y are divided by four, MNCS leaves overall performance unchanged, whereas CPP/FCSm decreases group A’s value.The subfield-specific performance of both groups remains unchanged under this rescaling.
4. Example 2
Example 2 shows that CPP/FCSm and MNCS recommend different investments: CPP/FCSm favors physics, whereas MNCS favors chemistry. This difference reflects CPP/FCSm’s bias toward fields with higher expected citation counts.
- Investment scenarios: The faculty must invest limited funds in equipment for either chemistry or physics, with each scenario expected to improve average publication performance.Scenario 1 invests in chemists, while scenario 2 invests in physicists.
- Investment scenarios: Because physics has an expected citation rate twice as high as chemistry, investment in chemistry appears preferable on substantive grounds.The comparison concerns the expected number of citations per publication in the two fields.
- Indicator-based decisions: The CPP/FCSm indicator instead directs the investment toward physicists, despite the apparent preference for chemistry.This conclusion is based on the expected effect on the faculty’s overall CPP/FCSm value.
- Indicator-based decisions: The MNCS indicator directs the investment toward chemists, which the example presents as the better decision given the available information.The result follows from comparing the expected effects on the faculty’s overall MNCS value.
- Indicator bias: CPP/FCSm favors physics because its calculation overweights publications in fields with higher expected citation counts.In this example, physics publications are overweighted relative to chemistry publications.
5. Consistency of indicators
This section examines consistency for total- and average-performance indicators and shows that MNCS uniquely combines consistency with homogeneous normalization. TNCS is consistent, whereas brute force and CPP/FCSm are not.
- Consistency definitions: Consistency is defined separately for indicators of total and average performance.The section introduces mathematical notation and formal definitions for both types of consistency.
- Total performance: TNCS is consistent, but the brute force indicator is not.Adding the same publication to two sets can change the brute force ranking, demonstrating inconsistency.
- Average performance: MNCS is consistent, whereas CPP/FCSm is not.In the example, adding the same publication causes CPP/FCSm to decrease from 3 to 1 for one set.
- Homogeneous normalization: Both CPP/FCSm and MNCS satisfy homogeneous normalization.For publications from one field, this property requires the indicator to equal average citations divided by the field’s expected citations per publication.
- Uniqueness theorem: MNCS is the only average-performance indicator satisfying both homogeneous normalization and consistency.Consequently, MNCS is the only direct alternative to CPP/FCSm that has consistency; other consistent indicators lack homogeneous normalization.
6. How to handle overlapping fields?
For overlapping fields, the MNCS indicator should assign publications proportionally across fields and use harmonic rather than arithmetic averaging. This approach preserves the desired value of one for all publications across all fields.
- How to handle overlapping fields?: The section examines how to calculate MNCS when Web of Science subject categories overlap and publications belong to multiple fields.The desired property is an MNCS value of one for the complete set of publications across all fields.
- How to handle overlapping fields?: In the example, publication 5 belongs to both fields X and Y, so it receives a weight of 1/2 in field-specific calculations.The example contains three fields and five publications; publication 5 is the overlapping publication.
- How to handle overlapping fields?: Using the arithmetic average of publication 5’s expected citations in fields X and Y violates the desired property because the resulting MNCS is not one.The arithmetic-average approach gives e5 = 5 and is described as unsatisfactory.
- How to handle overlapping fields?: Using harmonic averages yields an MNCS of one for all publications across all fields and is therefore the most appropriate treatment of overlapping fields.The same result can be obtained by calculating publication 5’s expected citations as the harmonic average of its field-specific expected citations.
7. Discussion and conclusion
The discussion contrasts the CPP/FCSm and MNCS normalization mechanisms, noting counterintuitive CPP/FCSm results and CWTS’s move toward MNCS while stressing MNCS’s limitations and complementary use.
- 7. Discussion and conclusion: The CPP/FCSm indicator sometimes yields counterintuitive results, prompting CWTS to move toward MNCS as its new crown indicator.The paper cautions that MNCS should not be used in isolation, but together with other indicators.
- 7. Discussion and conclusion: The MNCS indicator is an arithmetic average, so skewed citation distributions can make one highly cited publication dominate its value.Confidence intervals can provide additional information about this instability.
- 7. Discussion and conclusion: Very recent publications can strongly influence MNCS because their expected citation counts are often close to zero.Even one or two citations can therefore produce a large ratio of actual to expected citations.
- 7. Discussion and conclusion: MNCS requires a field-classification scheme, but fuzzy disciplinary boundaries and multidisciplinary research make such classifications partly arbitrary.A source-normalized indicator is suggested as a complementary approach.
- 7. Discussion and conclusion: MNCS gives publications equal weight across fields, a principle viewed as more reasonable than CPP/FCSm’s emphasis on high-citation fields.However, equal weighting may not always be completely satisfactory.
Appendix
The appendix proves that any bibliometric indicator satisfying homogeneous normalization and consistency must equal the MNCS indicator. The proof derives an aggregation identity and uses mathematical induction to establish equality with the MNCS definition.
- Theorem 1 proof: Theorem 1 assumes that fA has homogeneous normalization and consistency, then aims to show that fA is the MNCS indicator defined in (2).These are the two stated properties used throughout the proof.
- Theorem 1 proof: Homogeneous normalization and consistency are combined to derive equation (16) for all S ∈ Σ and all (c, e) ∈ P.The proof rescales S and a singleton publication using rational representations before applying consistency.
- Theorem 1 proof: Mathematical induction based on equation (16) establishes equation (21) for every finite multiset S = {(c1, e1), …, (cn, en)} ∈ Σ.The induction extends the aggregation identity from the constructed cases to arbitrary publication multisets.
- Theorem 1 proof: Equation (21) equals equation (2) by homogeneous normalization, so fA is the MNCS indicator and the theorem is proved.The appendix explicitly concludes the proof after identifying fA with the MNCS definition.