Source-linked AI summary
A critical comparative analysis of five world university rankings
Henk F. Moed
TL;DR
World university rankings claim to measure academic excellence, but their consistency and interpretability are limited by differences in coverage, indicators, calculations, and normalization. The paper compares ARWU, Leiden, THE, QS, and U-Multirank across these dimensions and finds partial alignment alongside meaningful divergence, arguing for tools that expose patterns across multifaceted data rather than only finalized indicator values.
Problem
The five ranking systems claim to measure academic excellence, but their consistency and the implications of their differences for stand-alone use require assessment.
Method
The paper links five ranking systems at the institutional level and compares their coverage, geographical distributions, indicator calculations, skewness, correlations, and selected indicator combinations.
Results
Spearman correlations between seven citation-, reputation-, or teaching-related indicators range from 0.4 to 0.8 for most pairs, indicating related but complementary measures rather than a single operationalization of excellence.
Takeaways & Limitations
Ranking positions reflect each system’s institutional coverage, indicator choices, rating methods, and normalizations, so users need insight into these orientations when interpreting results.
Takeaways & Limitations
Institutional comparisons are constrained when ranking entries do not clearly identify which university-system components or campuses they cover.
Abstract
from arXiv · showhide
To provide users insight into the value and limits of world university rankings, a comparative analysis is conducted of 5 ranking systems: ARWU, Leiden, THE, QS and U-Multirank. It links these systems with one another at the level of individual institutions, and analyses the overlap in institutional coverage, geographical coverage, how indicators are calculated from raw data, the skewness of indicator distributions, and statistical correlations between indicators. Four secondary analyses are presented investigating national academic systems and selected pairs of indicators. It is argued that current systems are still one-dimensional in the sense that they provide finalized, seemingly unrelated indicator values rather than offering a data set and tools to observe patterns in multi-faceted data. By systematically comparing different systems, more insight is provided into how their institutional coverage, rating methods, the selection of indicators and their normalizations influence the ranking positions of given institutions.
1. Introduction
The paper examines whether five university ranking systems consistently measure academic excellence and how their coverage, indicators, calculations, and normalizations shape interpretation. It proposes a comparative analysis to help users understand and use rankings more responsibly.
- Motivation: The study is motivated by growing demand for systematic evaluation of government-supported research and by the influence and continuing debate surrounding university rankings.Governments use evaluations to inform research allocation, support, organizational restructuring, and productivity initiatives.
- Research focus: The study compares ARWU, Leiden, QS, THE, and U-Multirank as systems that claim to provide information about academic excellence.ARWU, THE, and QS publish composite overall indicators, whereas Leiden and U-Multirank do not.
- Research focus: Its central question is whether the systems are consistent and, if not, which differences explain their divergent results.The paper also considers implications for interpreting and using any single ranking as a stand-alone information source.
- Analytical design: The first part analyzes institutional overlap, geographical coverage, indicator distributions and skewness, and statistical correlations between indicators.These analyses examine both the systems’ coverage and how their indicators represent academic excellence.
- Analytical design: The second part combines indicators, especially across systems, in four analyses intended to generate more detailed and comprehensive views of what indicators measure.One analysis examines correlations between citation- and reputation-based indicators in major national academic systems.
2. Analysis of institutional overlap
The overlap analysis matches institutions across five ranking systems and finds limited agreement in both overall coverage and top-100 membership. Interpretation is complicated by uncertainty over which campuses or components different entries represent.
- Institutional matching: 1,715 unique institutions and 3,248 variant names were identified; 377 universities (22 per cent) appear in all 5 systems, and 182 (11 per cent) in 4 systems.The matching process included manual inspection, name variants, and checks of institutions appearing in one system’s top 100 but not others.
- Matching limitation: Institutional matching is limited by ambiguous coverage of university systems and campuses, so unclear cases were treated as different institutions even when substantial overlap was possible.The problem is illustrated by differing entries for the University of Arkansas System across ARWU, Leiden, QS, THE, and U-Multirank.
- Top-100 overlap: 194 institutions occur in the combined top-100 lists, but only 35 appear in all five lists.For the top-100 comparison, ARWU, QS, and THE use overall rankings, while Leiden uses publication- and citation-based lists and U-Multirank has no overall ranking.
- Top-100 overlap: Top-100 overlap ranges from 49 institutions between the two Leiden lists to 75 between QS and THE.The Leiden lists differ by whether they are size-dependent or size-independent.
3. Geographical distributions
The ranking systems differ in their geographical orientation and in which institutions appear uniquely in their top-100 lists. These differences indicate that institutional and top-list coverage varies across systems.
- Country preferences: The preference ratio compares each country’s observed institutional representation in a ranking with its expected representation under independence.A value of 1.0 indicates that a country’s representation is as expected.
- Country preferences: U-Multirank is oriented towards Europe, ARWU towards North America and Western Europe, Leiden towards emerging Asian countries and North America, and QS and THE towards Anglo-Saxon countries.
- Unique institutions: The top-100 analysis identifies unique institutions as universities appearing in one system’s top list but not in any other system’s top list.
- Unique institutions: Most unique institutions in the ARWU and Leiden CIT top lists are from the USA, while QS unique institutions are concentrated in Great Britain, Korea and Hong Kong.
- Unique institutions: Leiden PUB unique institutions are especially located in China and Italy, whereas THE unique institutions are concentrated in Germany, the USA and The Netherlands.
4. Indicator scores and their distributions
The five systems transform raw indicators through different normalization and scoring methods, while data availability limits coverage. Their indicator distributions vary substantially in skewness and in the proportions assigned to U-Multirank performance groups.
- Coverage and normalization: ARWU, THE and QS publish overall indicators only for their first 100, 200 and 400 universities, respectively.QS also publishes all indicator values only for its first 400 institutions, and occasional values are missing.
- Coverage and normalization: ARWU and QS normalize each indicator by assigning the highest-scoring institution 100 and expressing other scores as percentages of that maximum.QS may apply cut-offs so multiple institutions receive a score of 100.
- Coverage and normalization: THE uses percentile ranks, cumulative probability functions and Z-scoring for most indicators, with an added exponential component for the Academic Reputation Survey.Figure 1 plots THE indicator scores against percentile-rank scores; citation observations lie on the diagonal.
- Coverage and normalization: U-Multirank assigns institutions to performance groups A–E according to distance from the median, producing distributions that can differ strongly from quintiles.For absolute publications, the A–E shares are 2.6%, 47.3%, 25.5%, 20.7% and 0.0%, respectively; for publications cited in patents they are 30.6%, 7.4%, 11.6%, 30.3% and 8.8%.
- Indicator distributions: Leiden’s absolute number of top publications has the highest skewness, whereas THE citations has the lowest; excluding them, ARWU indicators are most skewed and QS indicators are among the least skewed.
5. Statistical correlations
Correlations between ostensibly similar indicators vary widely across systems, while citation-impact measures from Leiden are closely aligned. Reputation and teaching measures can also correlate strongly when they share major survey components.
- Similar indicators: The QS Faculty-Student Ratio correlates only moderately with THE’s student-staff ratio (rho=-0.47), while QS International Faculty correlates very weakly with U-Multirank International Academic Staff.
- Citation indicators: The two Leiden citation-impact measures correlate very strongly (rho=0.98), whether impact is measured by mean normalized citation score or the top of the citation distribution.
- Citation indicators: THE’s Citation indicator also correlates strongly with Leiden impact measures, despite using Scopus while Leiden uses Web of Science.
- Reputation and teaching: THE Research and THE Teaching show the only very strong correlation among the citation-, reputation- and teaching-related indicators.Both are composite indicators in which reputation-survey outcomes form the major component.
- Citation indicators: QS Citations per Faculty correlates weakly with other citation indicators and has a weak within-QS correlation with academic reputation (Rho=0.34).
6. Secondary analyses
The secondary analyses compare indicator relationships across countries and ranking systems, showing that normalization choices and indicator construction can produce divergent institutional positions. They also identify interpretive limits, including unclear causes of national correlation patterns and possible affiliation effects in ARWU’s highly cited researcher measure.
- National academic systems: Italy, Brazil, and Russia contain universities with similar Research Performance but widely varying citation scores, whereas the Netherlands and Germany show the reverse pattern.Both patterns produce low rank correlation coefficients.
- Citation indicators: QS Citations Per Faculty and Leiden top-10-percent publications are compared as percentile ranks for institutions in six countries.The analysis examines whether the QS measure’s regional normalization affects cross-country comparability.
- Citation indicators: A second QS normalization produces negative correlations with the Leiden measure for Italy, the Netherlands, and especially Germany.Humboldt University Berlin and the University of Heidelberg have Leiden percentile ranks above 60 but QS Citation per Faculty percentile ranks below 20.
- Reputation indicators: Differences between THE Research Performance and QS Academic Reputation are probably linked to QS regional response-rate weightings, which THE does not apply.The comparison identifies institutions with the largest and smallest differences between the two percentile-rank measures.
- Highly cited researchers: Two Saudi Arabian institutions have much higher ARWU Highly Cited Researchers scores than expected from their Leiden top-publications output.The analysis connects this discrepancy to possible secondary affiliations in the Thomson Reuters highly cited researcher data, while noting that the interpretation requires further investigation.
- Highly cited researchers: The ARWU highly cited researcher indicator combines counts from 2001 and 2013 lists, while the newer list uses only authors’ primary affiliations.This substantially reduces the secondary-affiliation effect highlighted in the comparison.
7. Discussion and conclusions
Comparing five ranking systems shows that institutional coverage, geographical orientation, indicator construction, normalization, and data quality substantially shape rankings. The analysis therefore favors transparent, multidimensional tools over treating any system as a definitive measure of excellence.
- Institutional coverage: Only 35 institutions appear in all five systems’ top 100 lists, while pairwise overlap ranges from 49 to 75 institutions.The identity of the top 100 therefore depends on which ranking system is used.
- Geographical coverage: The systems define the world differently, with U-Multirank oriented toward Europe, ARWU toward North America, Leiden toward emerging Asia, and QS and THE toward Anglo-Saxon countries.
- Indicator construction: Four methods convert raw data into scores: maximum normalization, percentile ranks with some exponential components, and distance to the median.These choices affect score interpretation; in THE, 90 per cent of institutions score below 55 for Research Performance and below 50 for Teaching Performance.
- Indicator distributions: ARWU and THE indicators are generally more skewed than QS and Leiden indicators, showing that distributional properties differ substantially across systems.ARWU shows the largest skewness, followed by THE; QS and the two Leiden relative citation indicators show the lowest skewness.
- Indicator relations: Citation indicators from Leiden, THE, ARWU, and U-Multirank correlate strongly or very strongly, whereas QS Citation per Faculty correlates only weakly with them.The QS measure divides field-corrected total citations by faculty numbers and may also normalize by geographical region, although further research is needed.
- Indicator relations: Most pairwise correlations among seven citation-, reputation-, and teaching-related indicators are moderate or strong, with Spearman coefficients between 0.4 and 0.8, but never very strong.The findings support complementarity among systems and no single perfect operationalization of academic excellence.
- Interpretation and transparency: Ranking positions can be influenced by normalization and weighting choices, so users need simple tools showing their actual effects.The paper presents two-dimensional scatterplots as examples of tools for viewing multidimensional institutional data.
- Interpretation and transparency: Comparing indicator pairs can reveal discrepancies that warrant further study and help evaluate data quality and indicator validity.The paper also argues that systems should expose empirical data and tools for examining such patterns rather than only finalized rankings.
THE World University
THE covers multiple performance dimensions through weighted indicators, including teaching, research, citations, industry income, and international outlook. The supplied passages also identify institutional participation and a broader set of teaching, learning, and engagement measures.
- Coverage: The passages state that about 1,300 institutions are included and that, in principle, all higher education institutions can register for participation.
- Performance indicators: THE assigns 30% each to Teaching and Citations, 20% to Publications, 10% to International Faculty, 2.5% to Industry Income, and 20% to International Students.
- Performance indicators: The listed dimensions include teaching and learning, research, knowledge transfer, international orientation, and regional engagement.Examples include survey-based teaching quality, citation rate, regional income, and spin-offs.
- Indicator components: One supplied passage states that THE indicators are based on reputation, while another lists citations, industry income, and international outlook among its components.