Source-linked AI summary
Evaluating Journal Quality: A Review of Journal Citation Indicators and Ranking in Business and Management
John Mingers, Liying Yang
TL;DR
Journal indicators are increasingly used to evaluate research performance, but their differing designs and biases complicate journal-quality comparisons. The paper reviews these indicators theoretically and tests them empirically on business and management journals. Although the indicators appear highly correlated, they can produce substantially different rankings, so no single indicator is superior; h-index and SNIP are suggested as useful options.
Problem
Research-performance evaluation increasingly depends on journal indicators, but their differing designs, validity, and biases require systematic comparison.
Method
The paper analyzes indicators’ explicit and implicit biases theoretically and tests them empirically on business and management journals.
Results
Highly correlated indicators nevertheless produce considerable differences in individual journal rankings.
Takeaways & Limitations
No single indicator is superior; the paper suggests h-index and SNIP as potentially effective metrics.
Takeaways & Limitations
The study was conducted only within one disciplinary area.
Abstract
from arXiv · showhide
Evaluating the quality of academic journal is becoming increasing important within the context of research performance evaluation. Traditionally, journals have been ranked by peer review lists such as that of the Association of Business Schools (UK) or though their journal impact factor (JIF). However, several new indicators have been developed, such as the h-index, SJR, SNIP and the Eigenfactor which take into account different factors and therefore have their own particular biases. In this paper we evaluate these metrics both theoretically and also through an empirical study of a large set of business and management journals. We show that even though the indicators appear highly correlated in fact they lead to large differences in journal rankings. We contextualize our results in terms of the UK's large scale research assessment exercise (the RAE/REF) and particularly the ABS journal ranking list. We conclude that no one indicator is superior but that the h-index (which includes the productivity of a journal) and SNIP (which aims to normalize for field effects) may be the most effective at the moment.
1. INTRODUCTION
Research-performance evaluation increasingly relies on journal quality as a proxy, making accurate, robust, transparent, and unbiased indicators important. This paper reviews indicator biases and tests them empirically against business and management journals.
- Motivation: Research-performance evaluations increasingly use journal quality as a proxy for the quality of published papers.This practice is especially consequential in large-scale assessments such as the UK REF, which affect funding and careers.
- Existing approaches: Journal quality can be assessed through peer review, citation indicators, or hybrid approaches such as the ABS list.The ABS list combines peer-review ranking with citation indicators.
- Evaluation criteria: The growing use of journal indicators makes accuracy, robustness, transparency, and unbiasedness essential for confident evaluation.The paper emphasizes these criteria for users, particularly non-bibliometricians.
- Research gap: New indicators—including the Eigenfactor, h-index, SJR, and SNIP—require comparison because each differs from the JIF and may embody distinct biases.The paper examines how these indicators differ theoretically and how valid they are in practice.
- Approach: The paper analyzes indicator theories and their explicit and implicit biases, then tests them on business and management journals.The study also compares empirical results with the ABS list and develops recommendations.
2. REVIEW OF JOURNAL CITATION INDICATORS
The review examines citation-data sources and journal indicators, emphasizing coverage, data-quality, field-normalization, productivity, and citation-distribution differences. These design choices affect how journals are compared across fields and publication profiles.
- Citation-data sources: Google Scholar covers more research outputs than specialized databases, including books and reports, with coverage usually around 90%.Its coverage is comparatively advantageous in non-science subjects.
- Citation-data sources: Google Scholar can generate two to five times as many citations for a work because it draws from a wider range of sources.Its broader coverage is partly responsible for the larger citation counts.
- Citation-data sources: Google Scholar data quality is poor because entries may be duplicated and citations may come from varied non-research sources.The review also notes normalization difficulties and unresolved errors in specialized databases.
- Normalization: Citation practices vary substantially across fields, making cross-field journal comparisons difficult without field or source normalization.Differences in citation density and author counts contribute to this problem.
- Basic indicators: The JIF is a two-year mean of citations in year t to papers published during the previous two years.The review also questions its transparency and reports that its calculation could not be reproduced in one study.
- Basic indicators: The h-index combines citation impact and journal productivity but favors journals publishing many papers over those with fewer highly cited papers.It is limited by total publication volume and can disadvantage highly cited journals with relatively few publications.
- Second-generation indicators: SNIP normalizes for both publication volume and field, using reference journals specific to each journal.This design distinguishes it from measures that normalize only publication volume or neither dimension.
3. Methodology and Data
The study compares journal indicators using business and management journals drawn from multiple databases, then examines their correlations and ranking relationships. Database coverage and availability constrain comparability, while the indicators show strong but non-identical associations.
- Data and design: The study empirically compares journal indicators across a sample of business and management journals and relates the results to the ABS ranking.
- Data limitations: The databases do not cover the same journals, producing differences in citation counts and limiting consistency across indicators.
- Data and design: Data were collected for 2012 and 2013 from Web of Science, Scopus, and Scimago, with title and ISSN consistency checks and some outlier removal.
- Descriptive statistics: The dataset contains citation metrics with widely different scales, and most variables are highly skewed; the immediacy index is the main non-citation exception.
- Correlation analysis: Pearson correlations were calculated after inspecting indicator relationships, with impact-factor variants and several citations-per-paper measures showing very high correlations.
- Correlation analysis: High correlations do not establish identical measurement: prior evidence shows substantial ranking differences even when Eigenfactor and total citations correlate at 0.995.
4. Analysis of the Results
Principal-components and journal-ranking analyses show that indicator differences visible in theory also appear empirically. Indicators cluster according to normalization, citation volume, and prestige, while individual journal rankings can diverge substantially.
- Principal components: Principal-components analysis identifies aggregate patterns linked to whether indicators normalize citation counts and account for prestige.
- Principal components: PC2 separates indicators by paper normalization, with unnormalized measures taking positive values.
- Principal components: The main cluster includes JIF, 5-JIF, IPP, and SNIP, which normalize citations by the number of papers.
- Indicator relationships: SJR is closer to non-prestige indicators than to AIS, while SNIP is close to impact factors, suggesting limited visible effects of prestige and source normalization in this sample.
- Principal components: PC2 and PC3 clearly distinguish prestige-based indicators—SJR, Eigenfactor, and AIS—from indicators that do not incorporate prestige.
- Principal components: The indicators form four groups combining citation volume or citations per paper with prestige or its absence.
- Journal rankings: The summed-rank table favors journals that perform consistently across indicators, so journals doing very well on most measures but poorly on one or two may be excluded.
SNIP Rank
SNIP and related indicators produce materially different journal rankings because they emphasize citations per paper, normalization, or prestige in different ways. The results show substantial rank shifts and support using multiple indicators rather than treating any single measure as definitive.
- Rank differences: Even among top journals, the h-index can rank some titles over 50 places lower, while Eigenfactor and IPP produce similar effects.
- Rank differences: SNIP and IPP can rank Long Range Planning 30 places higher, while Brookings Papers is ranked 1st on article influence score but 111th on h-index.
- Indicator effects: The h-index is strongly affected by publication volume: higher-ranked journals average 267 documents per year versus 27 for lower-ranked journals.
- Indicator effects: Citation-per-paper indicators favor journals with highly cited, relatively few papers and disadvantage journals publishing more papers, even when those papers are highly cited.
- SNIP: SNIP aims to normalize citation density and publication volume, but the sample provides limited evidence that it substantially normalizes for field.
- Journal contrasts: MIS Quarterly ranks highly on impact factors and SNIP but lower on SJR and AIS, whereas finance journals rank higher on prestige measures than on pure citation measures.
- Overall implications: The indicators are highly correlated overall, yet individual journals can shift by over 100 places between measures, potentially changing classifications from 4* to 2* in ABS terms.
- Prestige and transparency: Prestige indicators partly distinguish journals, but their effects may reflect field differences; SJR and SNIP are also difficult to investigate because of limited transparency and reproducibility.
5. Comparing Journal Indicators with Peer Review Journal Lists
The section compares peer-reviewed journal lists, especially the ABS list, with citation-indicator rankings and finds systematic discrepancies between them. These differences are concentrated across fields and raise questions about the justification for relying on either approach alone.
- Peer-reviewed lists and their criticisms: Peer-reviewed lists remain the predominant practical basis for journal ranking, despite guidelines urging appropriate use of metrics.The UK ABS list is predominant despite intense criticism.
- Peer-reviewed lists and their criticisms: The ABS list is criticized for favoring traditional US-operated positivistic journals over eclectic and innovative European and non-US journals.This criticism specifically concerns the composition of ABS 4* journals.
- Indicator rankings versus ABS: Indicator rankings identify journals with high scores but low ABS grades, including two operations-management, five information-systems/management and two operations-research journals.Technovation is rated 2* and the Journal of Supply Chain Management only 1* in ABS.
- Field-level comparison: Field-level ABS and indicator rankings have a rank correlation of 0.61, while information management and operations management perform poorly in ABS and business history, accounting and finance relatively well.The pattern agrees with other research reporting that some fields are undervalued relative to their journals’ citation performance.
- Overall conclusion: Overall, the ABS list and citation-indicator rankings show systematic discrepancies, questioning the justification for relying on either ranking basis without comparison.The section emphasizes that such lists affect universities, departments and individual scholars despite extensive criticism.
6. Conclusions
The indicators appear highly correlated, but theoretical and empirical differences produce substantially different journal rankings. The authors therefore recommend cautious, pluralistic use of metrics and peer review, especially when interpreting journal-level results for individuals.
- Limitations and future research: Citation distributions are highly skewed, and using journal-level characteristics to judge individual papers or researchers risks an ecological fallacy.The paper therefore cautions against inferring individual performance from general journal-level results.
- Empirical findings: Highly correlated indicators can still produce significant differences in individual-journal rankings.These differences arise from distinct assumptions about normalization, field effects, citing-journal prestige, skewness, data quality, databases, and reproducibility.
- Empirical findings: Citation metrics and an established journal list often disagreed: highly cited journals could rank relatively low, while highly ranked journals could have little citation impact.The comparison was made against a journal list used extensively in research assessments.
- Recommendations: No single indicator is superior, so ranked lists should combine several metrics with peer review and be interpreted cautiously.The authors also note that peer review and expert journal lists are subjective and biased in many ways.
- Recommendations: SNIP and the h-index are the authors’ preferred metrics, because SNIP normalizes for papers and field while the h-index is transparent, understandable, and robust to poor data.The authors present these as recommendations rather than as universally superior measures.
- Limitations and future research: The study is limited to business and management journals and to the journals that could be included given data-source constraints.The authors call for larger tests of normalization, improved Google Scholar data, aggregated indices, and additional indicators.