Source-linked AI summary

Characterizing the Google Books corpus: Strong limits to inferences of socio-cultural and linguistic evolution

Eitan Adam Pechenick, Christopher M. Danforth, Peter Sheridan Dodds

arXiv:1501.00960v4physics.soc-phcond-mat.stat-mechcs.CLstat.AP

TL;DR

The paper examines how limitations of the Google Books corpus constrain interpretations of word and phrase frequency trends. It uses statistical divergence and individual-word contributions to characterize changes across English data sets, highlighting scientific and medical journals as an important challenge while retaining n-gram evolution as a lens on language and culture.

  • Problem

    Google Books frequency trends cannot directly measure true cultural popularity because the corpus is library-like and includes scientific and medical journals.

  • Method

    The paper compares statistical divergence between years and analyzes the contributions of individual words, using Jensen-Shannon divergence when vocabulary differs across years.

  • Results

    Scientific and technical language, punctuation shifts, historical terms, and author- or work-specific words contribute substantially to divergence across the examined English data sets.

  • Takeaways & Limitations

    N-gram evolution remains a valuable lens into changes in language use and culture, but Google Books should be characterized carefully before drawing broad conclusions.

  • Takeaways & Limitations

    Even after filtering scientific terms, normalized frequencies remain affected by print delays and prolific authors and cannot directly measure true cultural popularity.

Abstract

from arXiv · show

It is tempting to treat frequency trends from the Google Books data sets as indicators of the "true" popularity of various words and phrases. Doing so allows us to draw quantitatively strong conclusions about the evolution of cultural perception of a given topic, such as time or gender. However, the Google Books corpus suffers from a number of limitations which make it an obscure mask of cultural popularity. A primary issue is that the corpus is in effect a library, containing one of each book. A single, prolific author is thereby able to noticeably insert new phrases into the Google Books lexicon, whether the author is widely read or not. With this understood, the Google Books corpus remains an important data set to be considered more lexicon-like than text-like. Here, we show that a distinct problematic feature arises from the inclusion of scientific texts, which have become an increasingly substantive portion of the corpus throughout the 1900s. The result is a surge of phrases typical to academic articles but less common in general, such as references to time in the form of citations. We highlight these dynamics by examining and comparing major contributions to the statistical divergence of English data sets between decades in the period 1800--2000. We find that only the English Fiction data set from the second version of the corpus is not heavily affected by professional texts, in clear contrast to the first version of the fiction data set and both unfiltered English data sets. Our findings emphasize the need to fully characterize the dynamics of the Google Books corpus before using these data sets to draw broad conclusions about cultural and linguistic evolution.

I. INTRODUCTION

The Google Books corpus is valuable for studying changes in word use, but its library-like sampling and growing scientific content constrain cultural and linguistic inferences. The paper develops a principled way to examine word evolution and shows why corpus dynamics require careful characterization.

  • Corpus and analytical value: Google Books contains millions of digitized books represented as case-sensitive n-grams, making it exceptionally large but more lexicon-like than text-like.The first version contains over 5 million books and over half a trillion words; the second contains 8 million books.
  • Corpus and analytical value: Because the corpus contains roughly one copy of each book, n-gram frequencies reflect library presence rather than readers’ large-scale cultural popularity.Authors, editors, and publishers shape which books enter the library, while new editions and reprints can create additional appearances.
  • Corpus limitations: Scientific and medical texts increasingly influence the corpus, especially the unfiltered English sets and the first English Fiction version.The rise of capitalized “Figure” relative to “figure” provides a suggestive indication of scientific-text sampling from around 1900 onward.
  • Corpus dynamics: The corpus volume generally grows exponentially over time but dips during major conflicts, while English volume rises between versions and English Fiction volume decreases drastically.The fiction decrease suggests more rigorous filtering in the later version.
  • Corpus limitations: Individual authors can disproportionately affect corpus trends because authors are represented roughly by prolificacy rather than popularity.The paper highlights this issue through examples such as “Frodo” and the repeated appearance of Lanny Budd across 11 novels.

II. METHODS

The paper compares yearly 1-gram distributions using Jensen-Shannon divergence, a bounded and symmetric alternative to KL divergence. JSD avoids infinite divergence when a word occurs in one year but not the other and supports an information-theoretic interpretation.

  • Statistical divergence between years: The analysis calculates statistical divergence between 1-gram distributions from two years using Jensen-Shannon divergence.The distributions P and Q describe word probabilities in the first and second years.
  • Statistical divergence between years: KL divergence can become infinite when a 1-gram has positive probability in one year but zero probability in the other.Because this situation is common in the data sets, the analysis uses JSD instead.
  • Statistical divergence between years: Jensen-Shannon divergence is bounded from 0 for identical distributions to 1 bit when the distributions have no 1-gram overlap.It is also symmetric with respect to the two distributions.
  • Statistical divergence between years: KL divergence measures the average number of wasted encoding bits when text from one year is encoded using the other year’s distribution.The mixed distribution in JSD takes the place of the approximating distribution regardless of the text’s year.

B. Key contributions of individual words

The paper decomposes Jensen-Shannon divergence into contributions from individual words and examines how frequency changes, thresholds, and time comparisons shape those contributions. Common words can matter through subtle changes, while rare or new words can matter through sufficiently large shifts.

  • Key contributions of individual words: Each word’s contribution is proportional to its average probability, with the proportionality determined by the ratio between its smaller probability and average.This decomposition enables sorting and analyzing individual-word contributions between time periods.
  • Key contributions of individual words: Words with larger average probability or smaller probability ratios contribute more strongly to Jensen-Shannon divergence.Thus, both frequent words changing subtly and uncommon words undergoing large shifts can contribute substantially.
  • Key contributions of individual words: A word contributes at most its average probability in bits, reached exactly when its smaller probability is zero; unchanged probabilities contribute nothing.The contribution function is symmetric around r = 1, representing no change.
  • Key contributions of individual words: The analysis averages each word’s normalized frequency across equally weighted years within each decade before calculating and sorting divergence contributions.This coarse-grains comparisons such as 1800–1809 versus 1990–1999.
  • Key contributions of individual words: The heatmaps compare JSD for every year pair from 1800 to 2000 using words above normalized frequency 10^-5, with color scaled to each data set’s maximum divergence.The maps also mark comparisons involving 1880 and show year-to-year divergences off the diagonal.
  • Key contributions of individual words: The 1880 comparisons generally show an initial jump followed by a more gradual divergence increase, with unusually large wartime divergences that later return toward the prior trend.The figure uses a 10^-5 threshold for all words and 10^-4 for dashed curves.

A. Broad view of language evolution within Google Books

Jensen-Shannon divergence shows that Google’s underlying lexicon generally changes more as periods become farther apart, with wartime and recent-decade patterns standing out. Word-shift comparisons reveal contributions from war-related language, technical terms, citations, and prolific authors.

  • Broad temporal dynamics: The JSD heatmaps show divergence generally increasing with temporal distance, although this growth is strongly curtailed in the second English Fiction corpus.The heatmaps also show wartime pinching toward the diagonal and greater divergence between earlier and recent periods.
  • Broad temporal dynamics: Consecutive-year divergences typically decline through the mid-19th century, remain relatively steady until the mid-20th century, and then decline gradually.Recent decades show a more persistent boost in divergence than wartime bumps.
  • Inter-decade word shifts: The 1940s show increased references to Hitler, war, and other World War II-related military and political terms across the examined data sets.Word-shift figures compare the 1930s with the 1940s using the top individual 1-gram contributions to JSD.
  • Author effects: In fiction data sets, the recurring character Lanny Budd contributes noticeably to divergence, illustrating the influence of a prolific author.Lanny and Budd increase in relative frequency, and the character dominates the charts in the fiction data sets.

2. Second unfiltered English data set: 1930s versus 1940s

In the second unfiltered English data set, the 1930s-to-1940s divergence is dominated by increased references to years. War-related words also rise, while personal pronouns and “King” decline in relative frequency.

  • 2. Second unfiltered English data set: 1930s versus 1940s: Eight of the top ten contributions are years from 1940 through 1949, whose contributions decrease chronologically.“1948” and “1949” appear at ranks 15 and 34, respectively.
  • 2. Second unfiltered English data set: 1930s versus 1940s: “War” ranks 11th and increases in relative frequency between the 1930s and 1940s.
  • 2. Second unfiltered English data set: 1930s versus 1940s: “Hitler” and “Nazi” increase in relative frequency, ranking 18th and 26th, respectively.
  • 2. Second unfiltered English data set: 1930s versus 1940s: Parentheses rank 13th and 14th and show increased relative frequencies of use.
  • 2. Second unfiltered English data set: 1930s versus 1940s: Personal pronouns decrease in relative frequency, while “King” ranks 41st with a decreased relative frequency.The passage suggests the decrease in “King” may relate to the British line of succession.

3. Second unfiltered English data set: 1950s versus 1980s

Between the 1950s and 1980s, the second unfiltered English data set shifts toward punctuation, years, and technical vocabulary. These contributions suggest that computational and professional textual sources strongly shape the observed divergence.

  • 3. Second unfiltered English data set: 1950s versus 1980s: The top two contributions are parentheses, which show dramatically increased relative frequencies of use.
  • 3. Second unfiltered English data set: 1950s versus 1980s: Technical terms including “model,” “data,” “percent,” “Figure,” “technology,” and “information” show noticeable increases.
  • 3. Second unfiltered English data set: 1950s versus 1980s: 19 out of the top 30 contributions are increased relative frequencies for years between 1968 and 1980.
  • 3. Second unfiltered English data set: 1950s versus 1980s: “The,” “of,” and “which” decrease noticeably, while masculine pronouns decrease and “women” increases within the top 60.

4. First English fiction data set: 1930s versus 1940s

The first English Fiction data set resembles unfiltered English in its decade-to-decade word shifts, with years, technical and medical terms, and punctuation contributing prominently. Its divergence also reflects recurring fictional characters and prolific authors, so the label does not indicate primarily fictional content.

  • 4. First English fiction data set: 1930s versus 1940s: “Lanny” rises from rank 49th to 8th, while “King” leaves the top 60 and “patient” enters at rank 51st.
  • 5. First English fiction data set: 1950s versus 1980s: The medical and citation-like word shifts demonstrate that the first fiction data set is strongly influenced by medical journals and cannot be considered primarily fiction despite its label.
  • 6. Second English fiction data set: 1930s versus 1940s: In the second English Fiction data set, quotation marks and “Lanny” are the two greatest contributions between the 1930s and 1940s.“Budd” ranks 11th ahead of “Hitler” at 13th, and Lanny receives more mentions than Hitler during this period.
  • 6. Second English fiction data set: 1930s versus 1940s: The second fiction data set’s prominent contributions include numerous fictional characters from works by authors such as James T. Farrell, John Galsworthy, Pearl S. Buck, and Kathryn Forbes.Examples include Studs Lonigan, Dinny Cherrel, Wang Yuan, and “Mama.”
  • 6. Second English fiction data set: 1930s versus 1940s: The greatest contributions to divergence appear to correspond to the most prolific authors, particularly Upton Sinclair.The passages note possible non-fictional references for Marcel Proust and B. M. Bower, so not every prominent word is unambiguously fictional.

7. Second English fiction data set: 1930s versus 1940s

The second English Fiction data set shows varied shifts in pronouns, contractions, modality, profanity, honorifics, punctuation, and technical-term usage across decades. Technical terms remain relatively stable in fiction, except for the growing use of “computer.”

  • Masculine pronouns such as “he” and “himself” decrease, while feminine pronouns including “her,” “she,” and “She” increase.
  • Relative frequencies of contractions increase, while “shall” and “must” decrease and profanity increases.
  • Technical terms occur far less frequently and remain relatively stable in fiction, except for “computer,” which gains popularity since the 1960s.
  • “Mr.” and “Mrs.” decrease, alongside punctuation shifts including fewer semicolons and more periods.
  • Quotation and question marks increase in the 1980s, while the four-period ellipsis loses ground to the three-period version.

C. The rise of scientific literature in the Google Books corpus

Scientific and professional texts increasingly shape the unfiltered Google Books data, producing technical vocabulary and rising year-reference peaks that are largely absent from the second fiction data set. The comparisons also show that divergence reflects sampled journals, OCR changes, historical events, and prolific authorship.

  • Unfiltered English data sets feature general scientific terms, while the original fiction data set additionally includes medical terms among major divergence contributions.
  • Year-reference peaks rise in unfiltered data but not in fiction, strongly suggesting that scientific-journal citation bias drives the rise in the larger data set.
  • The declining half-lives of year mentions may reflect scientific citation dynamics in which older papers are cited less often amid expanding literature.
  • In the second fiction data set, “computer” gains popularity while other technical words remain relatively steady.

IV. CONCLUDING REMARKS

The Google Books corpus is not an unbiased sample: scientific publications increasingly dominate it, and even the first fiction data set appears saturated with medical literature. The authors therefore call for separating popular and scientific components while treating the corpus as a limited proxy for cultural popularity.

  • The corpus does not represent an unbiased sample of publications, and scientific publications increasingly dominate its contents throughout the 1900s.
  • Even the first data set labeled as fiction appears saturated with medical literature.
  • Future analyses should identify and distinguish popular and scientific components before using the corpus to study cultural and linguistic evolution.
  • Filtering scientific terms cannot make normalized frequencies a direct measure of true cultural popularity because the corpus remains library-like and is affected by author prolificacy and print delays.
  • The authors recommend focusing on the second English Fiction data set for English analyses or properly accounting for biases in unfiltered data.
Loading 1501.00960v4…