Source-linked AI summary
The Multidimensional Assessment of Scholarly Research Impact
Henk F. Moed, Gali Halevi
TL;DR
The article further develops multidimensional research assessment and considers how metrics inform assessment design and policy decisions. It shows that publication counts or journal metrics may be unsuitable for comparative performance assessment, without being invalid in every circumstance.
Problem
Publication counts have little value for comparing research-active scientists according to their research performance.
Method
The article further develops multidimensional research assessment and examines patent analysis and meta-analysis as assessment and policy tools.
Results
Assessment metrics can inform decisions about an assessment process’s overall objective and general setup, while journal metrics or publication counts may not suit comparative assessment.
Takeaways & Limitations
Metric use should be considered in relation to the assessment process and its policy decisions.
Takeaways & Limitations
The use of journal metrics or publication counts is not invalid under all circumstances.
Abstract
from arXiv · showhide
This article introduces the Multidimensional Research Assessment Matrix of scientific output. Its base notion holds that the choice of metrics to be applied in a research assessment process depends upon the unit of assessment, the research dimension to be assessed, and the purposes and policy context of the assessment. An indicator may by highly useful within one assessment process, but less so in another. For instance, publication counts are useful tools to help discriminating between those staff members who are research active, and those who are not, but are of little value if active scientists are to be compared one another according to their research performance. This paper gives a systematic account of the potential usefulness and limitations of a set of 10 important metrics including altmetrics, applied at the level of individual articles, individual researchers, research groups and institutions. It presents a typology of research impact dimensions, and indicates which metrics are the most appropriate to measure each dimension. It introduces the concept of a meta-analysis of the units under assessment in which metrics are not used as tools to evaluate individual units, but to reach policy inferences regarding the objectives and general setup of an assessment process.
Section 1: Introduction
The paper argues that research assessment is increasingly important to stakeholders and should match its methods and metrics to the assessment purpose, objectives, policy context, unit, and impact dimension. It develops a multidimensional assessment matrix covering common and newer indicators, impact dimensions, and policy-level analysis.
- Assessment framework: Assessment methods and metrics should be selected according to the purpose, objectives, and policy context of the evaluation.The framework treats assessment as a combination of methods that can be modeled, changed, and combined over time in line with stated goals.
- Motivation: Research assessment is a key concern for institutions, policymakers, administrators, program directors, and researchers seeking funding, promotion, or tenure.The paper links this concern to research quality, competition for talent and budgets, transparency, and accountability.
- Scope and contribution: The article systematically examines the usefulness and limitations of 10 frequently used metrics across articles, researchers, groups, departments, networks, and institutions.The set includes newer indicators such as the H-Index, usage indicators based on full-text downloads, and altmetrics based on social-media mentions.
- Scope and contribution: The paper presents an extended typology of research-impact dimensions and identifies which metrics are most appropriate for measuring each dimension.Its focus is on research impact, including scientific, economic, educational, and cultural aspects.
- Assessment framework: Publication counts can distinguish research-active staff from inactive staff but have little value for comparing active scientists’ research performance.The example illustrates why an indicator’s usefulness depends on the assessment task rather than being universally valid.
- Meta-analysis: The paper introduces a meta-analysis in which metrics inform policy decisions about an assessment process’s overall objective and general setup rather than evaluating individual units.This shifts the use of metrics from unit-level judgment toward policy conclusions about the assessment design.
Section 2: Main types of metrics and analytical tools
The section surveys major research-assessment metrics and analytical tools, emphasizing their uses across assessment levels and their methodological limitations. It covers publication and citation indicators, usage data, and emerging altmetrics.
- Usage and altmetrics: Altmetrics extend assessment through mentions in blogs, websites, Twitter, and other social-media channels.The section presents altmetrics as an emerging area accompanying traditional bibliometric approaches.
- Publication- and citation-based indicators: Publication and citation indicators remain central tools for assessing individual, institutional, program, and country-level scientific impact.Citation analysis can measure publication, author, institution or department, and country impact.
- Limitations: Large citation datasets can reduce random error, but systematic bias may affect particular subgroups and prevent errors from canceling out.The text cautions that sufficiently large samples do not automatically eliminate systematic deviations.
- Limitations: Citation counts may omit relevant work, reflect self-citations or closed communities, and fail to capture all literature read or influential.The section discusses uncited foundational work, self-citation, consensus effects, and “obliteration by incorporation.”
- Limitations: Citation-based comparisons are distorted by disciplinary, publication-age, document-type, database-coverage, and naming differences.The paper notes that normalized indicators can account for some field differences, whereas absolute counts may be distorted.
- Usage and altmetrics: Usage data such as clickstreams, downloads, and views can complement citations because some documents reach large audiences without being cited often.Reviews, editorials, tutorials, and technical outputs may be read more than cited.
Section 3: Assessment models
The section describes assessment models that combine expert judgment, diverse indicators, and self- or external evaluation. Hybrid models adapt the combination of measures to the field and assessment target.
- Assessment models: Assessment methods can combine peer review with bibliometric, econometric, altmetric, and end-user indicators.The distinction is between expert judgment and metrics-based assessment using multiple indicator types.
- Assessment models: Evaluation models also distinguish self-evaluation from external evaluation by an outside agency.Self-evaluation is described as critical reflection on one’s own performance.
- Hybrid models: Hybrid models combine different measures and approaches according to the field and target of assessment.Their modular design responds to discipline-specific assessment challenges.
- Hybrid models: Using varied measures and methods is presented as important for capturing institutional, program, and individual impact and productivity.The section links comprehensive evaluation with assessment at multiple organizational levels.
Program Assessment - Empowerment Evaluation (EE)
The section presents Empowerment Evaluation alongside institutional and university-ranking models. These approaches involve stakeholders, compare organizations through multiple dimensions, or provide information for benchmarking rather than a single ranking.
- Empowerment Evaluation: Empowerment Evaluation emphasizes involvement of both evaluators and evaluated participants in the assessment process.The model is described as participant-driven and centered on stakeholder involvement.
- Empowerment Evaluation: Empowerment Evaluation makes stakeholders responsible for selecting appropriate methods and metrics and for using the resulting data to improve processes.Stakeholders participate in data collection, analysis, interpretation, and implementation of improvements.
- University-ranking models: University-ranking models use bibliometric and other indicators to compare institutions across scientific, professional, web, social, and economic dimensions.Examples include ARWU, Leiden, SCImago, Webometrics, and broader multidimensional rankings.
- University-ranking models: Some systems provide benchmarking information rather than ranking universities, while others compare institutions across several activities.GRBS supports comparison in traditional and interdisciplinary areas, and U-Multirank includes teaching, research, knowledge transfer, international orientation, and regional engagement.
- Limitations: The section notes that rankings are debated and that diverse indicator sets reflect agreement that no single indicator captures quality or impact.Ranking findings and their advantages and disadvantages remain contested.
Section 4: The multi-dimensional research assessment matrix
The multidimensional assessment matrix links indicator choice to the assessment unit, research dimension, objectives, and policy context. It organizes metrics by their uses and limitations across levels and impact dimensions.
- Selection principles: The matrix extends assessment beyond bibliometric indicators by considering non-bibliometric measures, aggregation levels, impact sub-dimensions, and policy objectives.Its framework develops the principle that metric usefulness varies with assessment circumstances.
- Indicator matrix: The paper reviews the potentialities and limitations of 10 frequently used indicators, including publication, citation, patent, and altmetric measures.Table 4 summarizes strong points and limitations for the indicator set.
- Indicator matrix: Publication and citation measures serve different purposes, but their interpretation depends on field, publication age, document type, volume, and normalization.Examples include citation counts, citations per article, citation percentiles, and normalized indicators.
- Indicator limitations: Journal metrics cannot substitute for the citation impact of individual papers, while the H-Index is biased toward senior researchers.The section also notes that the H-Index is relatively insensitive to highly cited outliers and uncited articles.
- Impact dimensions: Research impact is divided into scientific-scholarly and societal categories, with societal impact including technological, social, economic, environmental, and cultural aspects.The typology associates these dimensions with publication, citation, collaboration, patent, media, and social-media indicators.
- Policy context: A meta-analysis of assessment units supports policy inferences about assessment objectives and general design rather than evaluating individual units directly.The policy context is tied to assumptions about the state or condition of the units under assessment.
- Selection principles: Indicator selection depends on the assessment unit, the research aspect being assessed, and the policy context.The unit’s discipline and mission also influence which indicators and data should be used.
Section 5: Discussion and conclusions
The discussion distinguishes using metrics to assess individual units from using them to analyze assessment systems and shape policy. It emphasizes that metrics and peer review each have strengths, limitations, and acceptable error rates that must be judged in context.
- Meta-analysis and policy: A meta-analysis examines the state of assessment units to inform the overall objectives and setup of an assessment process.It can also assist policy conclusions and help determine where peer assessment is necessary.
- Metric use and limits: Journal metrics and publication counts may help distinguish research-active from inactive staff, but are not sufficiently justified for comparing regular international-journal publishers by performance.The authors reject using indicators merely because they are easy to calculate and readily available, while allowing that such indicators may remain valid under some circumstances.
- Monitoring and acceptable error: Assessment methodologies should be monitored for intended and unintended effects, with severe effects such as metric manipulation potentially prompting replacement of the method.The authors note that both metrics and peer review carry risks of invalid outcomes, and that the scholarly and policy communities must decide acceptable error rates.
- Two levels of assessment: Metrics serve two levels: assessing particular units and providing insight into the functionality of a research system as a whole.At the system level, they help draw general conclusions about the system’s state.
- Meta-analysis and policy: Bibliometric studies can guide peer-review planning by identifying specialized fields or groups requiring thorough assessment.This is presented as one way to combine quantitative approaches with peer review in broad national exercises.
- Statistical considerations: Positive correlations between indicators do not necessarily hold across all value ranges, weakening simple justifications for comparative researcher assessment.A study of 1,325 journals found a significant overall correlation between acceptance rates and five-year impact factors, but much lower, nonsignificant correlations within quartiles.