Source-linked AI summary

Can we use Google Scholar to identify highly-cited documents?

Alberto Martín-Martín, Enrique Orduna-Malea, Anne-Wil Harzing, Emilio Delgado López-Cózar

arXiv:1804.10439v1cs.DL

TL;DR

The paper asks whether Google Scholar can reliably identify highly-cited documents, a capability important for bibliometric research on influential scientific works. It tests this using generic queries and rank–citation correlations, finding that Google Scholar identifies highly-cited papers effectively despite limited effects from language and operational factors.

  • Problem

    Prior research had not established whether Google Scholar could identify highly-cited documents, which represent influential scientific works important for bibliometric analysis.

  • Method

    The study uses generic queries filtered only by publication year and calculates the correlation between search-result position and received citations to minimize query-related ASEO effects.

  • Results

    A significant and high citation–ranking correlation shows that Google Scholar can effectively identify highly-cited papers, while publication-date and version-related errors have only incidental effects.

  • Takeaways & Limitations

    Google Scholar can reliably identify the most highly-cited academic documents and serves as a useful complementary bibliometric tool because of its wide coverage.

  • Takeaways & Limitations

    Language and geographic web-domain effects mainly affect results near the bottom of the ranking, and language restriction is procedurally imperfect.

Abstract

from arXiv · show

The main objective of this paper is to empirically test whether the identification of highly-cited documents through Google Scholar is feasible and reliable. To this end, we carried out a longitudinal analysis (1950 to 2013), running a generic query (filtered only by year of publication) to minimise the effects of academic search engine optimisation. This gave us a final sample of 64,000 documents (1,000 per year). The strong correlation between a document's citations and its position in the search results (r= -0.67) led us to conclude that Google Scholar is able to identify highly-cited papers effectively. This, combined with Google Scholar's unique coverage (no restrictions on document type and source), makes the academic search engine an invaluable tool for bibliometric research relating to the identification of the most influential scientific documents. We find evidence, however, that Google Scholar ranks those documents whose language (or geographical web domain) matches with the user's interface language higher than could be expected based on citations. Nonetheless, this language effect and other factors related to the Google Scholar's operation, i.e. the proper identification of versions and the date of publication, only have an incidental impact. They do not compromise the ability of Google Scholar to identify the highly-cited papers.

1. Introduction

Google Scholar’s broad scholarly coverage has supported research evaluation, but its ability to identify highly-cited documents had not been empirically established. The study therefore examines whether ranking position reliably reflects citation counts despite search-engine and document-level influences.

  • Background: Google Scholar indexes scholarly literature across disciplines, document types, and languages, while also providing citation counts and document-version information.Its open, automated model contrasts with the more controlled design of traditional bibliographic databases.
  • Prior research: Prior research has examined Google Scholar’s usability, comparative search performance, bibliographic and bibliometric data quality, unique citations, and coverage over time.This work established a substantial knowledge base about Google Scholar but did not directly test its ability to identify highly-cited documents.
  • Research problem: WoS Core Collection and Scopus identify highly-cited papers by sorting retrieved documents by citation counts, whereas Google Scholar lacks this sorting feature and limits queries to 1000 results.These differences raise the question of whether highly-cited papers can be reliably identified through Google Scholar.
  • Research problem: Academic Search Engine Optimisation modifies scholarly literature to improve its crawling, indexing, and ranking position.Google Scholar rankings may also reflect query factors and document characteristics in addition to citations.
  • Study objectives: The study aims to verify whether Google Scholar can reliably identify highly-cited papers and whether citations are the primary ordering criterion for generic queries.Highly-cited documents represent influential scientific works whose identification supports analysis of influential authors, topics, methods, and sources.

2. Method

The study used generic Google Scholar queries filtered only by publication year to reduce query-related optimisation effects, then compared result rank with citation counts. It assembled and checked 64,000 records across 1950–2013 and analysed the relationship using Spearman correlation.

  • Query design: Generic queries were defined as blank searches filtered only by publication year, reducing bias from specific keywords and query-related academic search engine optimisation.The approach was intended to test whether citation counts were associated with result ordering under minimally specified queries.
  • Sampling: 64,000 results were collected from 64 annual queries spanning 1950 to 2013, with a maximum of 1000 documents retrieved per query.The 1950 start year reflected an increase in Google Scholar coverage compared with preceding years.
  • Reliability check: The searches were performed twice using different network-access conditions, confirming that both datasets contained the same records before further processing.The two runs served as a reliability check involving a WoS Core Collection-connected computer and a normal Internet connection.
  • Variables: The study extracted result rank, Google Scholar citation counts, number of versions, and publication date, while document language was checked manually when possible.Rank represented each document’s position in the search results; citation counts reflected Google Scholar’s counts at query time.
  • Analysis: Spearman correlation was calculated because citation data followed a skewed distribution.The analysis related citation counts to search-result rank position.

3. Results

Google Scholar rankings generally track citation counts, especially within the first 900 results, supporting identification of highly cited documents. However, ranking irregularities become more pronounced near the bottom of the results, and yearly citation thresholds vary substantially.

  • r= -0.67 overall correlation linked citation counts to Google Scholar result positions across 64,000 documents.The annual correlations were stronger and stable, averaging -0.895 with σ=0.025.
  • r= 0.97 correlation characterized the top 900 positions, whereas the last 100 positions showed only r= 0.61.The weaker lower-ranking relationship included highly cited documents appearing in very low positions.
  • Only 11 of 64 yearly most-cited documents ranked first, while 32.8% ranked among the top three.The most extreme case occurred in 1978, when the most-cited document ranked 917th.
  • 60.9% of yearly documents ranked 1000 were also the least-cited documents, making position 1000 a useful citation-threshold indicator.The threshold was around 50 citations in 1950–1960, above 200 between 1985 and 2000, and around 50 again in 2010.

4. Discussion

The discussion examines Google Scholar’s stability and possible ranking distortions from search dynamics, document versions, and publication dates. These factors affect individual rankings unevenly but generally do not prevent identifying highly cited documents.

  • Dynamic nature of the academic search engine: Identical queries can produce different results across locations or over time, so individual rankings should be interpreted from a general perspective.Documents may appear, disappear, or move within the results page.
  • Google Scholar malfunction 1: Rank position and number of versions: The number of versions has a low but significant association with rank position (r = -0.30; α < 0.01), with a slight effect in the top 100 positions.The pattern is dispersed and lacks a stable year-by-year pattern.
  • Google Scholar malfunction 1: Rank position and number of versions: Version aggregation does not seem to greatly affect whether documents enter the top 1,000 results, although its exact positional effects were not systematically established.Incorrect aggregation could omit or wrongly attribute citations.
  • Google Scholar malfunction 2: Rank position and publication date: Only 2 of 64,000 documents failed the internal publication-date consistency test, but Google Scholar and Web of Science dates matched for 96.7% of linked documents.Date mismatches were concentrated in recent years, reaching 30.49% in 2012 and 76% in 2013; many 2013 errors exceeded 20 years.
  • Google Scholar malfunction 2: Rank position and publication date: Publication-date errors can place documents in the wrong year but do not exclude them from generic results, so they do not significantly impair identification of highly cited documents.Google Scholar’s selection of the latest edition of a monograph appears to be the primary cause of some mismatches.

ASEO document factor: Rank position and language of publication

Google Scholar ranks English-language documents disproportionately highly, while non-English documents are concentrated near the bottom of yearly result lists. Interface language and geographic web domain therefore affect ranking positions beyond citation counts.

  • 99.5 documents per year in the first 100 positions were English, indicating that other languages were unusually rare there.
  • 34.2% was the annual average share of English documents in the last 100 positions.
  • The high presence of highly cited non-English documents near the bottom may explain weaker correlations between ranking position and citations.
  • Restricting searches to Anglophone domains might raise correlations but would bias identification toward English-language publications.

5. Conclusions

Google Scholar’s citation-based ranking can reliably identify highly cited academic documents, supporting its use as a complementary bibliometric tool. Language, geographic-domain, publication-date, and version-related factors affect results, but the authors report that they do not compromise this capability.

  • Language and primary-version geographic domain mainly reduce ranking accuracy among approximately the last 100 results.
  • Publication-date errors and version-related effects have only incidental impact and do not compromise Google Scholar’s ability to find highly cited documents.
  • Google Scholar can reliably identify the most highly cited academic documents and serves as a complementary bibliometric tool.
Loading 1804.10439v1…