Source-linked AI summary

Coverage of highly-cited documents in Google Scholar, Web of Science, and Scopus: a multidisciplinary comparison

Alberto Martín-Martín, Enrique Orduna-Malea, Emilio Delgado López-Cózar

arXiv:1804.09479v3cs.DL

TL;DR

The study addresses whether selective bibliographic databases adequately represent highly-cited documents across fields and whether source choice affects bibliometric indicators. It analyzes 2,515 Google Scholar Classic Papers from 2006, comparing their coverage and citation counts with Web of Science and Scopus. Web of Science and Scopus miss substantial shares of highly-cited documents in Humanities, Literature & Arts and Social Sciences, while citation-count rankings remain strongly correlated across sources.

  • Problem

    Selective journal-based databases may omit highly-cited documents published in books, reports, conference papers, or non-selected journals, especially in locally oriented and non-English research areas.

  • Method

    The study compares 2,515 English original-research documents from Google Scholar’s 2006 Classic Papers with Web of Science and Scopus coverage and citation counts across broad subject categories.

  • Results

    Web of Science misses 28.2% of highly-cited documents in Humanities, Literature & Arts and 17.5% in Social Sciences, while Google Scholar citation counts correlate .83-.99 with Web of Science and Scopus counts.

  • Takeaways & Limitations

    Using selective databases to calculate highly-cited-document indicators might produce biased assessments in poorly covered areas, whereas coverage and citation data appear similar across sources in several natural-science fields.

  • Takeaways & Limitations

    The study analyzes only English original research published in 2006 and does not examine highly-cited Web of Science or Scopus documents missing from Google Scholar.

Abstract

from arXiv · show

This study explores the extent to which bibliometric indicators based on counts of highly-cited documents could be affected by the choice of data source. The initial hypothesis is that databases that rely on journal selection criteria for their document coverage may not necessarily provide an accurate representation of highly-cited documents across all subject areas, while inclusive databases, which give each document the chance to stand on its own merits, might be better suited to identify highly-cited documents. To test this hypothesis, an analysis of 2,515 highly-cited documents published in 2006 that Google Scholar displays in its Classic Papers product is carried out at the level of broad subject categories, checking whether these documents are also covered in Web of Science and Scopus, and whether the citation counts offered by the different sources are similar. The results show that a large fraction of highly-cited documents in the Social Sciences and Humanities (8.6%-28.2%) are invisible to Web of Science and Scopus. In the Natural, Life, and Health Sciences the proportion of missing highly-cited documents in Web of Science and Scopus is much lower. Furthermore, in all areas, Spearman correlation coefficients of citation counts in Google Scholar, as compared to Web of Science and Scopus citation counts, are remarkably strong (.83-.99). The main conclusion is that the data about highly-cited documents available in the inclusive database Google Scholar does indeed reveal significant coverage deficiencies in Web of Science and Scopus in several areas of research. Therefore, using these selective databases to compute bibliometric indicators based on counts of highly-cited documents might produce biased assessments in poorly covered areas.

Introduction

The study examines whether database choice affects highly-cited-document indicators, contrasting selective journal-based coverage with inclusive document-level indexing. It asks whether Google Scholar reveals coverage gaps in Web of Science and Scopus across subject areas and whether citation rankings agree.

  • Database selection can change the value of bibliometric indicators for a given unit of analysis.
  • Selective databases may miss highly-cited books, reports, conference papers, and articles in non-selected journals, potentially biasing assessments.Inclusive databases theoretically allow each document to qualify on its own merits.
  • Web of Science and Scopus have poor coverage in locally oriented Social Sciences and Humanities research and show a bias against non-English publications.
  • Evidence indicates that highly-cited documents increasingly appear in non-elite journals, partly because web search makes relevant articles easier to find.
  • Google Scholar’s Classic Papers lists highly-cited 2006 documents across 8 broad and 252 specific subject categories, using the top 10 documents per subcategory.Included documents presented original research, were published in English, and had at least 20 citations by May 2017.
  • The study poses questions about documents missing from Web of Science and Scopus, citation-rank similarity across databases, and which source provides the most citations.

Methods

The authors extracted Google Scholar Classic Papers data and checked its coverage and citation counts against Web of Science and Scopus. They analyzed coverage proportions, citation-count correlations, and category-level citation averages.

  • 2,515 records were extracted from Google Scholar’s Classic Papers dataset, which lists the most-cited 2006 documents within subject subcategories.French Studies contained 5 qualifying documents rather than 10.
  • The extracted data included subject categories, document bibliographic information, author profiles, publication venues, and citation counts recorded in May 2017.
  • Coverage in Web of Science Core Collection and Scopus was checked for the 2,515 documents, primarily by manually searching document DOIs.
  • Missing coverage was classified by causes including uncovered sources, incomplete source indexing, later indexing starts, and formally unpublished documents.
  • June 2017 data collection was retained after later searches found 2 additional Web of Science documents and 7 additional Scopus documents, whose citation exposure had changed.
  • Spearman correlations compared Google Scholar citation counts with Web of Science and Scopus, while log-transformed averages and 95% confidence intervals supported category-level comparisons.
  • The raw data, R code, and analysis results were made openly available.

Results

Across 2,515 highly-cited documents identified by Google Scholar, Web of Science and Scopus differ substantially in coverage across subject categories, while overlapping citation counts correlate strongly. Google Scholar reports higher citation counts in every category, with statistically significant differences throughout.

  • Coverage: 208 documents (8.2%) were absent from Web of Science and 87 (3.4%) from Scopus; 219 were absent from both databases.The jointly missing documents included 175 journal articles, 40 conference papers, one report, and three preprints.
  • Coverage: 28.2% of Humanities, Literature & Arts documents and 17.5% of Social Sciences documents were missing from Web of Science, compared with 17.1% and 8.6% from Scopus.Web of Science also missed 11.6% of Engineering & Computer Science and 6.0% of Business, Economics & Management documents; Scopus missed 2.5% and 2.7%, respectively.
  • Coverage: 56% of documents missing from Web of Science and 49% missing from Scopus came from journals or conferences indexed only after 2006.Because these databases generally do not index earlier documents retrospectively, publications preceding later coverage remain absent.
  • Citation counts: Spearman correlations between Google Scholar and the other databases ranged from .83 to .99 for documents covered by both sources.Google Scholar–Scopus correlations were stronger than Google Scholar–Web of Science correlations in every subject category; the highest value was .99 in Chemical & Material Sciences.
  • Citation counts: Google Scholar citation counts exceeded Web of Science and Scopus counts in every subject category, with statistically significant differences in all categories.The gap was larger in Business, Economics & Management, Social Sciences, and Humanities, Literature & Arts, and smallest in Chemical & Material Sciences.
  • Citation counts: Scopus had higher average log-transformed citation counts than Web of Science, but the difference was statistically significant in only 4 of 8 subject categories.Even in those categories, the confidence intervals were very close.

Limitations

The study’s evidence is constrained by the Classic Papers dataset and by its one-directional coverage comparison. Its conservative sample limits negative conclusions, while missing documents still challenge selective databases’ suitability for highly-cited-document indicators.

  • Dataset scope: Classic Papers displays only the top 10 most-cited English-language original-research documents published in 2006 within each subcategory.The fixed top-10 rule ignores variation in yearly publication output across subcategories.
  • Interpretation boundary: The study’s sample is an extremely conservative representation of highly-cited documents.Therefore, finding no missing documents in Web of Science or Scopus is not conclusive evidence of broad coverage, especially in high-output subcategories.
  • Interpretation boundary: Missing documents in Web of Science or Scopus within this highly exclusive sample still question those databases’ suitability for highly-cited-document indicators in some areas.
  • Coverage comparison: The study does not test how many highly-cited Web of Science and Scopus documents are absent from Google Scholar.This opposite-direction coverage analysis is identified as requiring a separate study.

Discussion and conclusions

The study finds substantial coverage gaps in Web of Science and Scopus for highly cited documents in several fields, while Google Scholar citation counts correlate strongly with both selective databases. These findings support field-sensitive database choices, balancing coverage against practical metadata and extraction constraints.

  • Coverage deficiencies: 28.2% of highly-cited Humanities, Literature & Arts documents are absent from Web of Science, compared with 17.1% from Scopus.For Social Sciences, the corresponding missing shares are 17.5% and 8.6%.
  • Coverage deficiencies: Web of Science also misses highly-cited documents in Engineering & Computer Science and Business, Economics & Management.In Computer Science, the study attributes the gap to weaker coverage of conference proceedings, an important publication venue in the field.
  • Citation-count agreement: Spearman correlations between Google Scholar and Web of Science or Scopus citation counts range from .83 to .99 across eight subject categories.The strongest reported value is .99 in Chemical & Material Sciences, while .83 occurs for Business, Economics & Management in the Google Scholar–Web of Science comparison.
  • Citation-count agreement: Google Scholar provides significantly higher citation counts than Web of Science and Scopus in all eight subject areas.Differences are larger in Business, Economics & Management, Humanities, Literature & Arts, and Social Sciences, suggesting a larger document base in these fields.
  • Implications: Inclusive databases have better highly-cited-document coverage than Web of Science or Scopus in several fields, whereas all three show similar coverage in four other areas.The similar-coverage areas are Health & Medical Sciences, Physics & Mathematics, Life Sciences & Earth Sciences, and Chemical & Material Sciences.
  • Implications: Database suitability depends on each bibliometric analysis because Google Scholar combines useful coverage with limited metadata and difficult extraction.The paper highlights missing author affiliations and funding acknowledgements, as well as the lack of an API; such assessments can affect hiring, promotion, funding, and rankings.
Loading 1804.09479v3…