Source-linked AI summary
About the size of Google Scholar: playing the numbers
Enrique Orduña-Malea, Juan Manuel Ayllón, Alberto Martín-Martín, Emilio Delgado López-Cózar
TL;DR
The paper addresses how much academic content Google Scholar contains and what its search results omit. It applies four empirical estimation methods and places Google Scholar’s May 2014 size at about 160 million documents, while emphasizing substantial methodological uncertainty.
Problem
Google Scholar’s changing, opaque indexing makes its size and coverage difficult to measure, despite search engines’ aspiration to index current academic knowledge.
Method
The paper applies four empirical approaches: capture-recapture, comparative database estimates, direct queries, and absurd Boolean queries.
Results
About 160 million unique Google Scholar records is the paper’s sensible overall estimate, while individual methods produce disparate values.
Takeaways & Limitations
The methods return similar results overall but differ in validity, so Google Scholar’s size remains an estimate rather than a directly reported figure.
Takeaways & Limitations
The estimate is constrained by low institutional-repository indexation, exclusion of files over 5 MB, and the capture-recapture method’s non-equivalent overlap measure.
Abstract
from arXiv · showhide
The emergence of academic search engines (Google Scholar and Microsoft Academic Search essentially) has revived and increased the interest in the size of the academic web, since their aspiration is to index the entirety of current academic knowledge. The search engine functionality and human search patterns lead us to believe, sometimes, that what you see in the search engine's results page is all that really exists. And, even when this is not true, we wonder which information is missing and why. The main objective of this working paper is to calculate the size of Google Scholar at present (May 2014). To do this, we present, apply and discuss up to 4 empirical methods: Khabsa & Giles's method, an estimate based on empirical data, and estimates based on direct queries and absurd queries. The results, despite providing disparate values, place the estimated size of Google Scholar in about 160 million documents. However, the fact that all methods show great inconsistencies, limitations and uncertainties, makes us wonder why Google does not simply provide this information to the scientific community if the company really knows this figure.
1. INTRODUCTION
Academic search engines have renewed efforts to measure the academic Web, but their coverage and indexing practices make size and missing-information questions difficult to answer.
- 1. INTRODUCTION: The academic-Web size debate concerns coverage, methodological accuracy, and the extent to which produced knowledge is indexed, searchable, retrievable, and accessible.
- 1. INTRODUCTION: Google Scholar’s aspiration to index current academic knowledge can make search results appear more complete than they are, raising questions about missing information.
- 1. INTRODUCTION: Academic search engines complicate size measurement because their indexed content is not fully controlled and their data can change over time.Unlike traditional bibliographic databases, their size and evolution cannot be obtained reliably through a simple record query.
- 1. INTRODUCTION: Opaque information policies, particularly at Google Scholar, further hinder independent assessment of database size and coverage.
- 1. INTRODUCTION: Khabsa and Giles estimated 114 million English academic-Web documents and about 99.8 million in Google Scholar, prompting questions about the estimate and its coverage claim.
2. OBJECTIVES
The paper aims to estimate Google Scholar’s size in May 2014 and assess the strengths and weaknesses of multiple empirical estimation methods.
- 2. OBJECTIVES: The study’s main objective is to calculate the size of Google Scholar in May 2014.
- 2. OBJECTIVES: It seeks to explain and apply various empirical methods for estimating Google Scholar’s size.
- 2. OBJECTIVES: It also evaluates the strengths and weaknesses of the different estimation methods.
3. METHODS
The paper estimates Google Scholar’s size using four empirical approaches: capture-recapture, comparative database evidence, direct queries, and Boolean absurd queries.
- 3. METHODS: Four procedures estimate Google Scholar’s size: Khabsa and Giles’s method, empirical database comparisons, direct database queries, and absurd Boolean queries.
- 3. METHODS: Khabsa and Giles’s method samples 150 English academic papers, compares their incoming citations in Microsoft Academic Search and Google Scholar, and applies capture-recapture logic.The sample covers 15 fields, and the citation overlap is measured with the Jaccard similarity index.
- 3. METHODS: The capture-recapture analogy differs from the original Lincoln-Petersen indicators because database sizes are compared while overlap is measured through citing documents.
- 3. METHODS: The comparative method synthesizes prior studies by calculating Google Scholar-to-Web of Science document ratios and combining them with geometric means and medians.The synthesis restricts comparisons by database, unit of analysis, and document language or language treatment.
- 3. METHODS: Direct queries use Google Scholar hit-count estimates by year and across 1700–2013, while also separating records, citations, and patents.Comparable queries are performed for Microsoft Academic Search and Web of Science.
- 3. METHODS: Absurd queries combine a common term with a nonexistent-site exclusion to approximate retrieval of all Google Scholar records, using both global and year-by-year ranges.
4. RESULTS
Across four estimation procedures, the paper reports widely varying Google Scholar size estimates while identifying substantial methodological and coverage limitations. The results nevertheless indicate a scale of roughly 150–160 million documents, with database biases and query inconsistencies limiting confidence in any single figure.
- 4.2. Estimates from empirical data: The four procedures produce disparate Google Scholar size estimates, with the paper’s empirical-data estimate reaching 153.5 million documents.The authors present these estimates alongside discussion of their advantages, disadvantages, and shortcomings.
- 4.1. Khabsa and Giles’s method: Khabsa and Giles estimate 114 million English academic-Web documents and 99.3 million English documents in Google Scholar, but their figures may be biased by method, universe, data, and database choices.The paper specifically questions treating Google Scholar and Microsoft Academic Search as the whole academic-web universe.
- Biases in the sample by document type: Journal documents comprise 75% of Web of Science records, while books and book chapters comprise 1%, creating under-representation of disciplines using other communication formats.The paper notes that Google Scholar and Microsoft Academic Search have different document-type distributions, with journal articles averaging 65% of Google Scholar records.
- Biases in the sample by document type: Spanish university output includes under 40% journal articles, about 30% books and chapters, and around 20% conference communications, indicating that traditional databases do not represent all scientific output.The authors state that such databases may cover less than 30% of total institutional scientific output in the Spanish case.
- 4.3.1. Sectional query: Google Scholar’s temporal query returned only 596,000 documents for 1700–2013 and produced inconsistent results, including more records for 2000–2013 than for longer periods.Single-year queries appeared more plausible, returning 141,000 results for 1900 and 2,410,000 for 2000.
- 4.3.2. Longitudinal query: The summed article records from 1700–2013 yield 99.8 million Google Scholar records, including 59.8 million documents written in English.The paper presents comparative data for Web of Science, Microsoft Academic Search, and Google Scholar alongside this query result.
5. CONCLUSIONS
The paper compares four empirical approaches to estimating Google Scholar’s size and judges them inconsistent, yet convergent enough to suggest about 160 million unique records. It also argues that database composition and opaque information policies prevent a definitive count.
- The estimates are constrained because Method A compares databases with different characteristics and the underlying samples overrepresent English-language articles relative to Google Scholar’s broader document types.The discussion notes that patents and case law should also be included, while some Google Scholar materials may not cite sampled articles.
- Method B estimates 171 million records, while correcting Method C.2 for possible errors yields around 114 million records.Method B is described as rough and imprecise; Method C.2 relies on hit-count estimates with unknown rounding and uncertain reliability.
- Four methods produce broadly similar results except direct-query Method C, whose estimates are unexpectedly more distant and whose C.1 variant is discarded.Method D is similar to Methods A and B but differs significantly from Method C.
- About 160 million unique records is presented as the sensible Google Scholar estimate after allowing for roughly 10% internal errors.This estimate excludes different versions of the same record.
- Google Scholar’s opaque information policies force researchers to infer a figure that the company could presumably provide directly.The paper questions whether Google itself knows or will publish the database size.
Notes
The notes describe browser-based querying procedures and identify external resources and a Web of Science coverage boundary.
- Custom-range queries can be generated directly through the browser after an initial Google Scholar hit-count estimate.The procedure selects the required time span without adding keywords.
APPENDIX I. CATALOGUE OF EMPIRICAL WORK RELATED TO THE SIZE OF GOOGLE SCHOLAR
The appendix catalogs empirical comparisons of Google Scholar with other databases across disciplines, document types, samples, and citation measures. The listed studies report varied coverage, overlap, and citation-count results.
- Google Scholar is compared with WoS, Scopus, MAS, and specialized databases using samples ranging from journals and books to researchers and field-specific publications.The catalogue spans multiple disciplines and study designs.
- Journal coverage comparisons report 40,000 journals for GSM versus 19,708 in JCR and 10,677 in SJR.
- Citation overlap varies substantially across studies, including 44 shared GS-and-Scopus citations, 11 GS-and-WoS citations, and 9 citations shared by all three databases in one comparison.
- GS reports 2.5 times as many citations as WoS in one comparison of books’ impact.
- GS identifies 1,448 more citations than the WoS-and-Scopus union, corresponding to 53.0% more citations in one LIS-faculty study.
AUTHORS SAMPLE TYPE ANALYSIS GS UNIT GS/GSM/GSC % WoS % SCO %
The sample-type analysis breaks down citations by document category across Google Scholar, Web of Science, and Scopus. Google Scholar includes a wider range of document types than the comparison databases in the listed example.
- Google Scholar’s citation sample includes articles, books, conference papers, foreign-language documents, theses, reports, syllabi, presentations, blogs, and web pages.
- Articles account for 59.6% of the GS sample, compared with 99.7% for WoS and 83.8% for Scopus.
- The GS sample contains 318 books, 329 dissertations, and 281 foreign-language documents, whereas WoS lists none for those categories in the displayed comparison.
- A second breakdown lists 5,493 GS citations, including 42.45% from journals, compared with 2,023 and 2,301 total citations in the other columns.
NO. OF PUBS FOUND IN GS
The cited studies examine publication and citation coverage in Google Scholar across publication types, journals, disciplines, and languages. One reported result finds that only 11.3% of papers missed by Google Scholar were non-English.
- The reviewed studies compare Google Scholar and Web of Science citations by publication type.
- Prior analyses sampled journals, articles, and faculty publications to examine Google Scholar coverage and citations.The samples include communication journals, nursing journals, open-access ISI-indexed journals, and LIS faculty work.
- One cited study examines citation counts for articles collected from the SDJ in September 2010.
- 11.3% of the papers missed by Google Scholar were non-English, compared with an average of 20.5% non-English materials in Compendex.The cited comparison spans disciplinary ranges of 10.8%–28.8% in Compendex.