Source-linked AI summary

Evaluating research: from informed peer review to bibliometrics

Giovanni Abramo, Ciriaco Andrea D'Angelo

arXiv:1811.01765v1cs.DL

TL;DR

National research assessments increasingly require methods that support comparison and policy decisions, but the relative performance of peer review and bibliometrics across measurement criteria needs evaluation. The paper contrasts both approaches using accuracy, robustness, validity, functionality, time, and costs. It concludes that bibliometrics is preferable for natural and formal sciences because databases can cover publications more broadly and support more robust, valid, functional, cheaper, and faster assessments.

  • Problem

    National research assessments are expanding, while evidence about applying peer-review and bibliometric measurement systems across core assessment parameters remains limited.

  • Method

    The paper contrasts peer review and bibliometrics in national research assessments using accuracy, robustness, validity, functionality, time, and costs.

  • Results

    Bibliometrics is by far preferable to peer review for natural and formal sciences because it can evaluate all publications indexed in WoS and Scopus, improving robustness, validity, functionality, costs, and execution time.

  • Takeaways & Limitations

    National publication databases derived from WoS or Scopus would support better, cheaper, and more frequent assessments for the natural and formal sciences.

  • Takeaways & Limitations

    Bibliometric indicators cannot cover the entire range of research outputs and are limited to publications and conference proceedings.

Abstract

from arXiv · show

National research assessment exercises are becoming regular events in ever more countries. The present work contrasts the peer-review and bibliometrics approaches in the conduct of these exercises. The comparison is conducted in terms of the essential parameters of any measurement system: accuracy, robustness, validity, functionality, time and costs. Empirical evidence shows that for the natural and formal sciences, the bibliometric methodology is by far preferable to peer-review. Setting up national databases of publications by individual authors, derived from Web of Science or Scopus databases, would allow much better, cheaper and more frequent national research assessments.

1. Introduction

National research assessments are expanding as governments use them to inform funding and other public objectives. The paper contrasts peer review with bibliometrics and argues that bibliometrics is preferable for natural and formal sciences because it can evaluate broader outputs across key measurement parameters.

  • National research exercises increasingly assess universities and public institutions to inform funding, improve performance, reduce information asymmetry, and demonstrate public benefits.
  • Peer review evaluates institution-submitted research through appointed expert panels, generally emphasizing output quality.
  • Informed peer review combines expert judgment with citation information and other quantitative indicators, as in the UK REF.
  • The Australian ERA mainly uses bibliometrics in natural and formal sciences, evaluating outputs against citation benchmarks without peer review and using volume indicators for overall performance.
  • The paper compares peer review and bibliometrics using accuracy, robustness, validity, functionality, time, and costs.
  • Bibliometrics is judged by far preferable for natural and formal sciences because peer review cannot feasibly evaluate an entire national research output, limiting productivity measurement and other properties.

2. Accuracy

The paper finds that peer judgment and bibliometric indicators have competing strengths for individual outputs. However, peer review introduces subjectivity and reviewer-selection concerns, while bibliometrics is limited by coverage and citation-timing issues.

  • Peer evaluation is susceptible to subjective distortions in judging quality, selecting experts, and selecting outputs for assessment.
  • Bibliometric indicators apply only to publications and conference proceedings, not the entire range of research outputs.
  • Citation counts may not always reflect quality, although negative citations are rare and do not disrupt the analyses.
  • Citation analysis is less reliable for recent works because citations may require time to mature, and delayed recognition can affect mature works.
  • Combining peer review with standardized bibliometric indicators lets reviewers compare subjective judgment with quantitative evidence, leaving the approaches broadly balanced for individual products.
  • Reviewer selection is critical to evaluation effectiveness, and quantitative performance indicators may help identify qualified reviewers as well as assess outputs.

3. Robustness

Robustness means that institutional rankings should not depend strongly on how much output is evaluated. Evidence from Italian disciplines shows that peer-review rankings can vary substantially with subset size, whereas larger subsets only partly stabilize results.

  • A robust assessment produces rankings that are not sensitive to the share of research output evaluated.
  • Changing the evaluated share from 25% of FTE researchers to 9% of total output changed rankings for 40 of 50 Physics universities, with shifts up to 15 positions.
  • In Biological sciences, the same subset change altered rankings for 45 of 53 universities, with a maximum shift of 22 positions.
  • Across eight Physics subset sizes ranging from 4.6% to 60% of total output, only 8 of 50 universities remained in the same ranking decile.
  • Physics rankings increasingly converge toward rankings based on all WoS-indexed publications, with only marginal correlation gains beyond the 30% scenario.
  • The authors conclude that peer review is intrinsically lacking in robustness, even when the evaluated subset is enlarged.

4. Validity

Validity is compromised when peer review assesses selected outputs rather than an institution’s complete research production. Internal selection can distort representation of quality through subjective choices, conflicts of interest, and difficulties comparing products across periods or subfields.

  • 4. Validity: Validity means measuring what counts, but restricting assessment to selected outputs compromises that objective.The paper defines validity as a measurement system’s ability to measure what counts.
  • 4. Validity: Peer-review exercises measure submitted products, not necessarily the institution’s best research outputs.The REF further selects staff and outputs, risking distortion if institutions do not identify their best researchers.
  • 4. Validity: Subset selection can introduce distortion because internal decisions may favor or block researchers or products rather than reflect intrinsic quality.The paper also identifies technical difficulties in comparing outputs from different periods and research subfields.
  • 4. Validity: The Italian VTR investigation found that selected outputs often did not represent institutions’ best products, so resulting rankings did not reflect their real quality.The authors attribute this to inefficiency in the selection process.

5. Functionality

Functionality concerns whether an assessment system serves all the policy purposes for which it is used. The paper argues that researcher- and group-level comparisons provide information needed for resource allocation and other objectives that peer review does not consistently supply.

  • 5. Functionality: Functionality is a measurement system’s ability to serve all the functions for which it is used.National assessments pursue objectives including efficient resource allocation.
  • 5. Functionality: Peer-review assessments do not consistently provide precise, comparable information for universities to allocate resources among individual researchers or research groups.Internal allocation may also be affected by personal interests that conflict with the institution’s collective interest.
  • 5. Functionality: Assessments comparing individual researchers and research groups would better support internal resource allocation, performance stimulation, and reduced information asymmetry.The paper states that bibliometrics has the capacity to provide this level of evaluation.

6. Cost and Time Effectiveness

Peer-review assessments require substantial direct and indirect costs and typically take years to implement, producing infrequent evaluation cycles. ORP-based bibliometrics is presented as substantially cheaper and faster.

  • 6. Cost and Time Effectiveness: £12 million was the direct cost of the UK’s 2008 RAE, while indirect institutional costs were estimated at five times that amount.The upcoming Italian VQR’s direct costs were estimated at €11 million.
  • 6. Cost and Time Effectiveness: Two years or more are needed to implement peer-review exercises, which typically recur every five to six years.The latest RAE covered an eight-year period, making these assessments slow and infrequent.
  • 6. Cost and Time Effectiveness: ORP’s direct costs are estimated at around 10% of peer-review exercises’ direct costs, with execution requiring only a few months.The shorter timeline and lower cost could support more frequent evaluations.

7. ORP and the bibliometric approach

The paper argues that bibliometrics is preferable for natural and formal sciences despite proxy limitations, especially when national databases directly identify researchers’ publications. The ORP demonstrates a large-scale implementation that supports comprehensive, researcher-level assessment with lower administrative burden.

  • 7. ORP and the bibliometric approach: Bibliometrics is presented as preferable to peer review for natural and formal sciences across robustness, validity, functionality, cost, and time effectiveness.The authors acknowledge that publications and citation counts remain proxies for total research output and publication quality.
  • 7. ORP and the bibliometric approach: Bibliometrics can evaluate all research output rather than a selected subset, avoiding distortions introduced by internal product selection.It also permits consideration of research quantity in addition to quality.
  • 7. ORP and the bibliometric approach: Australia’s bibliometric approach still requires institutions to submit selected products and maintain publication archives using researcher-provided data.That process retains direct and indirect opportunity costs for universities and researchers.
  • 7. ORP and the bibliometric approach: National database construction is difficult because institutional affiliations, author identities, homonyms, and name variations must be reconciled.These difficulties previously limited remote bibliometric evaluation to individual institutions or selected disciplines.
  • 7. ORP and the bibliometric approach: The ORP covers publications since 2001 from approximately 350 Italian public research organizations and attributes publications to academic authors with less than 5% error.It includes about 272,000 articles and reviews and 100,000 conference proceedings.
  • 7. ORP and the bibliometric approach: ORP-based indicators can assess individual researchers and aggregate results progressively to research groups and institutions.Indicators are standardized by subject-category citation intensity to limit differences in publication and citation practices.
  • 7. ORP and the bibliometric approach: The ORP decision-support system requires no institutional input, reducing indirect costs and enabling evaluations on the order of months rather than years.This non-invasive design is presented as enabling greater evaluation frequency.

8. Conclusions

Bibliometric evaluation is preferable to classic peer review for the natural and formal sciences because it can assess representative publication outputs more robustly, validly, functionally, cheaply, and quickly. The authors support national publication databases while retaining peer review for other disciplines.

  • Bibliometry can evaluate all publications indexed in WoS and Scopus, unlike peer review’s small evaluated subsets.The authors emphasize coverage of representative research output in the natural and formal sciences.
  • National databases such as Italy’s ORP attribute publications to precise authors with an acceptable 5% authorship error rate.For aggregate evaluation, sufficiently uniform false positives and negatives can limit ranking distortions.
  • Italy can use bibliometric methods for natural and formal sciences and institution-independent publication data instead of institutional submissions.The proposed approach combines bibliometric evaluation for these disciplines with peer review elsewhere.
  • Comparable national databases could enable international comparisons and more efficient decisions on incentives, recruitment, policy, and research strengths.The paper links comparative measurement to internal organizational systems and national or regional research and industrial policy.
  • The authors call for more work on large-scale measurement methodology and national databases, while recognizing peer review as more appropriate outside the natural and formal sciences.The conclusion limits the bibliometric recommendation by discipline.
Loading 1811.01765v1…