Source-linked AI summary

Large-scale comparison of bibliographic data sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic

Martijn Visser, Nees Jan van Eck, Ludo Waltman

arXiv:2005.10732v2cs.DL

TL;DR

The paper addresses limited large-scale evidence about how major bibliographic data sources differ in coverage and citation-link quality. It performs a document-level, pairwise comparison using Scopus as a baseline, finding that Microsoft Academic offers the most comprehensive coverage and arguing for comprehensive data with flexible filters. The analysis is limited by its reliance on datasets from 2018 and 2019.

  • Problem

    The study addresses the need to understand the strengths and weaknesses of bibliographic data sources for bibliometric research and practice.

  • Method

    The paper conducts a large-scale document-level comparison of five sources, using Scopus as a baseline for pairwise comparisons.

  • Results

    Microsoft Academic offers by far the most comprehensive coverage of the scientific literature among the five data sources studied.

  • Takeaways & Limitations

    Data sources should be as comprehensive as possible while providing filters that allow users to select literature according to their use case.

  • Takeaways & Limitations

    The analysis is not entirely up-to-date because it is based on datasets from 2018 and 2019.

Abstract

from arXiv · show

We present a large-scale comparison of five multidisciplinary bibliographic data sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic. The comparison considers scientific documents from the period 2008-2017 covered by these data sources. Scopus is compared in a pairwise manner with each of the other data sources. We first analyze differences between the data sources in the coverage of documents, focusing for instance on differences over time, differences per document type, and differences per discipline. We then study differences in the completeness and accuracy of citation links. Based on our analysis, we discuss strengths and weaknesses of the different data sources. We emphasize the importance of combining a comprehensive coverage of the scientific literature with a flexible set of filters for making selections of the literature.

1. Introduction

The paper addresses the need to understand the strengths and weaknesses of bibliographic data sources through a large-scale, document-level comparison. It compares Scopus pairwise with Web of Science, Dimensions, Crossref, and Microsoft Academic, examining document coverage and citation-link completeness and accuracy.

  • Motivation: Large-scale access constraints have made bibliographic data sources commonly compared through small case studies focused on fields, researchers, or limited document sets.Earlier comparisons also operated at journal level or used restricted document samples.
  • Contribution: The study compares five major multidisciplinary sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic.The comparison is conducted at the document level.
  • Contribution: The analysis examines differences in document coverage and in the completeness and accuracy of citation links.It considers scientific documents including journal articles, preprints, conference papers, books, and book chapters.
  • Scope: The comparison excludes Google Scholar, OpenCitations, and PubMed because of access, overlap, or citation-link constraints.Google Scholar lacks large-scale access, OpenCitations provides roughly the same data as Crossref, and PubMed lacks citation-link data.
  • Approach: Scopus serves as the baseline for pairwise comparisons with each of the other data sources because full Web of Science access is unavailable.Using Scopus as the baseline does not imply that it is the preferred source.

2. Data sources

The study uses source-specific datasets and scopes, comparing their differing content-selection policies and metadata practices. The sources range from selective collections to systems emphasizing comprehensive coverage, while Crossref depends on publisher-provided metadata.

  • Web of Science: Web of Science data come from the accessible Core Collection indices, excluding the Emerging Sources and Book Citation Indexes.The study labels this restricted dataset CWTS WoS rather than the full Web of Science database.
  • Selection policies: Dimensions and Microsoft Academic emphasize comprehensive coverage, whereas Web of Science emphasizes selectivity and Scopus appears more focused on database size.Dimensions also contains grants, datasets, clinical trials, patents, and policy documents, but these are not analyzed.
  • Crossref: Crossref makes publisher-supplied metadata openly available and its completeness and quality depend on what publishers provide.Crossref does not actively collect or enrich the data.

3. Matching of data sources

The paper matches Scopus documents to records in other sources using a staged, precision-oriented procedure based on normalized metadata and matching scores. The evaluation identifies high first-step coverage but also shows that incomplete or inconsistent metadata can prevent correct matches.

  • Preprocessing: The matching procedure preprocesses years, bibliographic numbers, author names, titles, source titles, and non-US-ASCII characters before identifying candidate pairs.Author initials and normalized text support comparisons across sources.
  • Scoring: Matches require a score above a threshold chosen to favor precision over recall, with only the highest-scoring match retained.The score compares DOI, author, title, source, year, volume, issue, pages, and article number; partial text matches use Levenshtein distance.
  • Matching criteria: Six consecutive matching steps move from restrictive criteria to less restrictive criteria using DOI, bibliographic fields, authors, source identifiers, and similar titles.The criteria include publication year, volume, pages or article numbers, and title-word overlap.
  • Efficiency: At least 80% of matches for each data source are made in the first matching step.Later steps handle fewer remaining documents, keeping computational cost acceptable.
  • Version handling: The one-to-one matching design can classify repository or alternate versions as unique content when only another version matches across sources.For example, a Scopus journal record may match a Microsoft Academic journal record while its repository version remains unmatched.

4. Comparison of coverage of documents

The five data sources differ substantially in size, overlap, document types, disciplines, languages, and temporal coverage. Microsoft Academic has the broadest coverage and largest overlap with Scopus, while CWTS WoS has the smallest overlap; high-level overlap statistics conceal important content differences.

  • CWTS WoS has the smallest overlap with Scopus, with almost 18 million shared documents, compared with 21 million for both Dimensions and Crossref.
  • Microsoft Academic includes some non-scientific content, but four of 30 sampled unmatched documents were clearly non-scientific, indicating a small share overall.
  • Dimensions and Crossref have similar yearly coverage, reflecting Dimensions’ strong reliance on Crossref data.
  • Coverage varies by document type: CWTS WoS misses meeting abstracts and book reviews, while Dimensions and Crossref include more book chapters and journal content absent from Scopus.
  • Matching is weaker for non-English documents: about 40% of Scopus records match Dimensions or Microsoft Academic, versus 21% matching CWTS WoS.

5. Comparison of completeness and accuracy of citation links

The comparison evaluates citation-link completeness and accuracy pairwise against Scopus, correcting for differences in document coverage. Scopus and CWTS WoS overlap most, while Crossref shows the largest discrepancy, largely because reference lists are unavailable or withheld.

  • Comparison method: Pairwise comparisons use only citation links between documents covered by both sources, separating citation-link quality from document-coverage differences.The analysis examines original citation links and excludes links recoverable only through alternative matching algorithms.
  • Scopus versus CWTS WoS: Scopus and CWTS WoS have the largest relative overlap, but 1.9% of CWTS WoS links are absent from Scopus and 5.8% of Scopus links are absent from CWTS WoS.These discrepancies may reflect incorrect identification or failure to identify links in either source.
  • Scopus versus Crossref: 57.9% of Scopus citation links cannot be obtained from Crossref, reflecting missing, undisclosed, or incorrectly unlinked reference lists.Crossref publishers may omit references, restrict their open availability, or be affected by technical linking problems.
  • Scopus versus Crossref: Dimensions provides access to many more citation links than Crossref despite their fairly similar document coverage, partly because it enriches Crossref data with direct publisher data.The enrichment includes citation links, abstracts, and affiliation data.
  • Citation-link accuracy: Manual examinations suggest high precision and recall for Scopus citation links, with few incorrectly identified links in Scopus.About half of sampled Dimensions links absent from Scopus were missed by Scopus, whereas only one sampled Microsoft Academic link absent from Scopus was missed by Scopus.
  • Citation-link accuracy: Microsoft Academic sometimes creates links for in-print references or secondary document versions where Scopus does not.These differing choices help explain some links present in Microsoft Academic but absent from Scopus.

6. Conclusions

The comparison finds distinct coverage and citation-link trade-offs across the five bibliographic data sources. It concludes that comprehensive coverage is best combined with flexible filters for selecting relevant literature.

  • Scopus and CWTS WoS: Scopus covers many documents absent from CWTS WoS, while nearly all CWTS WoS journal articles are also covered by Scopus.Scopus additionally covers many book chapters, whereas CWTS WoS covers meeting abstracts, book reviews, and some proceedings papers absent from Scopus.
  • Dimensions and Crossref: Dimensions and Crossref have similar document coverage and together cover many journal articles, book chapters, and proceedings papers not covered by Scopus.Many documents unique to Dimensions and Crossref are meeting abstracts or other short items, while Scopus also covers proceedings papers absent from them.
  • Coverage: Microsoft Academic offers the most comprehensive coverage, although its basic document-type classification provides limited insight into the nature of covered documents.A manual sample examination found that most additional documents were scientific and included many non-English documents.
  • Citation links: Scopus and CWTS WoS provide higher-quality citation links than Dimensions and Microsoft Academic, while all sources suffer from incomplete or inaccurate links.Missing citation links are especially a problem in Dimensions and Microsoft Academic, and CWTS WoS raises concerns about phantom citations.
  • Citation links: Crossref’s aggregate citation counts include closed citation links, but individual closed links cannot be accessed.This limits the accessibility of the underlying citation relations even when they contribute to aggregate counts.
  • Implications: An ideal data source combines comprehensive scientific-literature coverage with flexible filters for making selections suited to different purposes.The authors view comprehensiveness and selectivity as complementary rather than mutually exclusive.

7. Limitations

The analysis is based on data from 2018 and 2019, while the studied sources continue to change. Its conservative matching procedure likely underestimates overlap, and the comparisons exclude some direct source comparisons and two WoS citation indices.

  • Data from 2018 and 2019 means some findings may not represent the current state of the regularly improved and expanded data sources.The analysis is not entirely up-to-date, and recent developments are not covered.
  • The study compares Scopus pairwise with each other source, but does not directly compare CWTS WoS, Dimensions, Crossref, and Microsoft Academic with one another.
  • The Emerging Sources Citation Index and Book Citation Index are excluded from the CWTS WoS analysis.These indices are part of the WoS Core Collection, so the exclusion matters when interpreting findings for WoS.
  • The document-matching procedure prioritizes avoiding false positives over false negatives, so the reported overlap between Scopus and other sources is underestimated.

Competing interests

The authors are affiliated with CWTS at Leiden University, which has commercial relationships with the producers of several compared data sources. CWTS and Digital Science also collaborate through RoRI, while Waltman advises an open scholarly metadata research centre.

  • The authors are affiliated with the Centre for Science and Technology Studies at Leiden University.
  • CWTS has commercial relationships with Clarivate Analytics, Elsevier, and Digital Science, producers of Web of Science, Scopus, and Dimensions, respectively.
  • CWTS and Digital Science are founding partners of the Research on Research Institute.
  • Waltman serves as chair of the Advisory Board of the Research Centre for Open Scholarly Metadata and works with OpenCitations in that capacity.

Data availability

The study combines licensed, freely provided, and openly available data sources. Restrictions prevent redistribution of the Scopus, Web of Science, and Dimensions data, while the paper’s figure statistics are deposited in Zenodo.

  • Scopus and Dimensions data were made freely available to CWTS for research purposes, but cannot be redistributed.
  • Web of Science data were provided to CWTS under a paid license and cannot be made available.
  • Crossref data came from its openly available API, while Microsoft Academic data came from an openly available 2019 data dump.
  • The statistics presented in the paper’s figures are available in Zenodo.
Loading 2005.10732v2…