Source-linked AI summary
Big Macs and Eigenfactor Scores: Don't Let Correlation Coefficients Fool You
Jevin West, Theodore Bergstrom, Carl Bergstrom
TL;DR
The paper examines why high correlations between journal metrics can mislead comparisons of the information they provide. It argues that such correlations do not safely establish that metrics convey the same information.
Problem
The paper addresses whether highly correlated journal metrics necessarily provide the same information.
Method
The paper analyzes the relation between variables and compares journal metrics, including Eigenfactor, Impact Factor, and citation counts.
Results
High correlations, including values of 0.8, 0.9, or higher, do not safely show that two metrics convey the same information.
Takeaways & Limitations
Correlation coefficients should not be used alone to conclude that journal metrics provide equivalent information.
Takeaways & Limitations
The comparison between Eigenfactor and Total Citations may make little sense because Eigenfactor is a different type of measure.
Abstract
from arXiv · showhide
The Eigenfactor Metrics provide an alternative way of evaluating scholarly journals based on an iterative ranking procedure analogous to Google's PageRank algorithm. These metrics have recently been adopted by Thomson-Reuters and are listed alongside the Impact Factor in the Journal Citation Reports. But do these metrics differ sufficiently so as to be a useful addition to the bibliometric toolbox? Davis (2008) has argued otherwise, based on his finding of a 0.95 correlation coefficient between Eigenfactor score and Total Citations for a sample of journals in the field of medicine. This conclusion is mistaken; here we illustrate the basic statistical fallacy to which Davis succumbed. We provide a complete analysis of the 2006 Journal Citation Reports and demonstrate that there are statistically and economically significant differences between the information provided by the Eigenfactor Metrics and that provided by Impact Factor and Total Citations.
1 Big Macs and Correlation Coefficients
The Big Mac example shows that a near-perfect correlation can reflect shared currency-denomination variation rather than meaningful agreement between variables. Real wages, which measure purchasing power, still differ substantially across countries.
- Correlation example: 0.99 is the correlation between local-currency burger prices and hourly wages across 22 countries.The comparison uses Big Mac prices and mean hourly wages in local currency.
- Meaningful variable: Real wage, defined as burger prices divided by hourly wages, measures workers’ purchasing power.Its units are burgers per hour.
- Meaningful variable: Seven minutes in Denmark versus nearly two hours in China illustrates dramatically different purchasing power despite the high correlation.These values describe the work time required to afford a Big Mac at mean hourly wages.
- Why correlation misleads: Nominal hourly-wage variation is about 300 times larger than real-wage variation because currency denominations dominate the nominal measures.The paper nevertheless treats real-wage variation as more important for workers’ quality of life.
- Why correlation misleads: A high correlation does not imply constant real wages: the real-wage ratio’s standard deviation is 62% of its mean.The paper argues that differences in real wages are negligible only relative to much larger currency-denomination differences.
2 Davis’s analysis
Davis compared Eigenfactor metrics with citation-based measures and inferred that high correlations indicate similar information. The paper identifies statistical and measurement problems with that comparison and explains a more appropriate comparison.
- Davis’s comparison: Davis analyzed Eigenfactor scores and Total Citations for 165 medical journals, reporting a high correlation between them.He used a regression analysis of Eigenfactor scores on Total Citations.
- Davis’s comparison: Davis concluded that iterative citation weighting produced rankings not significantly different from raw citation counts.He characterized popularity and prestige measures as providing very similar information.
- Metric comparability: Eigenfactor scales with journal size, whereas Impact Factor measures citation impact per paper and should be independent of journal size.For a per-article comparison with Eigenfactor, the paper points to the Article Influence Score.
- Statistical critique: The paper identifies Davis’s comparison as a classic statistical error involving two measures with a common factor.It addresses this issue alongside the interpretation of correlation coefficients.
- Statistical critique: The paper rejects the inference that a high correlation coefficient means two alternative measures have no significant difference.It notes that Davis’s Eigenfactor–Impact Factor correlation was 0.86 but says the comparison makes little sense.
3 Journal Sizes and Spurious Correlations
Large variation in journal size creates spurious correlations when size enters both bibliometric measures, obscuring the distinction between popularity and prestige.
- Journal Sizes and Spurious Correlations: Journal sizes vary enormously, so correlations between size-dependent measures can reflect journal size rather than the relationship of interest.The JCR includes journals ranging from tiny publications to journals publishing tens of thousands of articles.
- Journal Sizes and Spurious Correlations: Davis’s regression effectively compares log(Article Influence) + log(Total Articles) with log(Impact Factor) + log(Total Articles).Total Articles acts as a common factor in both sides of the regression.
- Journal Sizes and Spurious Correlations: The shared log(Total Articles) term varies more than the other terms and obscures the relation between popularity and prestige.The intended comparison is between per-article measures rather than measures dominated by journal size.
- Journal Sizes and Spurious Correlations: Even uncorrelated Impact Factor and Article Influence could produce a high Eigenfactor–Total Citations correlation because both share number of articles as a common factor.The paper estimates a correlation of approximately 0.6 for all journals under this mechanism.
- Journal Sizes and Spurious Correlations: A high correlation between pages and total citations would not make pages an adequate surrogate for total citations.This analogy illustrates why journal-size-driven correlation does not establish equivalence between bibliometric measures.
- Journal Sizes and Spurious Correlations: 0.818 is the correlation for all 7,611 journals, while the mean across 231 fields is 0.853 with a standard deviation of 0.099.The medical sample’s ρ = 0.954 ranks in the 90th percentile among fields; within-field correlations typically exceed the aggregate correlation.
4 Correlation and significant differences
The paper tests whether Eigenfactor score is interchangeable with Total Citations by examining ratios rather than correlations, finding substantial and significant variation across journal groups.
- Correlation and significant differences: The relevant comparison is the Eigenfactor score-to-Total Citations ratio, because a shared journal-size factor cancels in the ratio.For A = ax and B = bx, A/B = a/b, so variance in x does not determine variance in the ratio.
- Correlation and significant differences: If Eigenfactor scores did not differ from Total Citation counts, EF/TC should be constant across journal groups.The analysis therefore examines normalized EF/TC ratios across journals and journal classifications.
- Correlation and significant differences: 1.42 × 10−5 is the mean EF/TC ratio for science journals versus 2.12 × 10−5 for social science journals.A Mann-Whitney U test finds this difference highly significant at p < 10−167.
- Correlation and significant differences: 49% is the difference in mean EF/TC ratios, implying that Total Citations undervalues social science journals relative to Eigenfactor-based valuation.The paper frames EF/TC as a measure of “bang per cite received.”
- Correlation and significant differences: 29% is the higher EF/TC ratio for public, environmental, and occupational health journals than for the rest of Davis’s medical sample.The difference is statistically significant, with Mann-Whitney U p < .01.
- Correlation and significant differences: The science–social science comparison is not unique; many other journal comparisons also yield significant EF/TC differences.This supports variation in the ratio beyond the single grouping examined in detail.
5 The value of visualization
The paper argues that visualizations reveal substantial ranking and valuation differences between bibliometric metrics that correlation summaries can obscure. Figures 3–5 expose both ordinal shifts and cardinal differences across journals.
- Figure 3: Figure 3 compares journal rankings by Total Citations and Eigenfactor score, using connecting lines to show upward, downward, or unchanged movement.
- Interpretation: The visual displays make ordinal differences in highly correlated data more apparent than a summary statistic such as the Spearman correlation.
- Figure 3: Aviation Space and Environmental Medicine drops 30 places while PLoS Medicine rises 31 places, contradicting the claim that journal ordering changes little.
- Figure 4: Figure 4 compares Impact Factor and Article Influence rankings for 84 journals, with ρ = 0.955 despite substantially different ordinal rankings.
- Figure 4: Top-ten journals change by only 1 or 2 positions, whereas larger rank changes occur farther down the hierarchy.
- Figure 5: Journal rankings can conceal cardinal differences: even journals retaining the same ordinal rank may receive substantially different valuations under two metrics.
6 Conclusion
The conclusion warns that correlation coefficients can mislead comparisons of bibliometric measures and recommends visualization and direct descriptive comparisons. The authors argue that Eigenfactor Metrics provide a substantially different view of journal prestige than raw citation counts.
- Correlation coefficients require considerably greater caution when spurious correlates are present or a formal hypothesis-testing framework is absent.The authors specifically caution against drawing conclusions from the magnitude of correlation coefficients alone.
- High correlation does not necessarily mean that two variables provide the same information, just as low correlation does not mean they are unrelated.The conclusion illustrates both cautions with wage and hamburger-price comparisons and successive logistic-map iterates.
- Comparison plots such as those in Figure 4 better highlight differences between bibliometric measures than standard scatter plots.The authors recommend choosing data graphics suited to the task because summary statistics can obscure important facets of the data.
- The median “bang per cite received” in the top third of journals is almost 2.4 times the median in the bottom third.The authors use this descriptive comparison to characterize differences in how journals are valued under the Eigenfactor Metrics.
- The Eigenfactor Metrics offer a substantially different view of journal prestige from straight citation counts.The conclusion links this difference to how journals are valued under the Eigenfactor Metrics.