Source-linked AI summary
Beyond word frequency: Bursts, lulls, and scaling in the temporal distributions of words
Eduardo G. Altmann, Janet B. Pierrehumbert, Adilson E. Motter
TL;DR
The paper investigates whether word recurrence dynamics exhibit scaling patterns beyond Zipf’s law. Using diverse text databases, it characterizes recurrence times and develops a generative model, finding stretched-exponential burstiness strongly related to semantic class and robust across datasets.
Problem
The paper asks whether successive word occurrences display statistical scaling regularities that are not captured by Zipf’s frequency law.
Method
The authors analyze recurrence times of frequent words across diverse databases and model word usage as a stationary renewal process whose occurrence probability depends on time since the previous occurrence.
Results
Word recurrence times are well described by a stretched exponential (Weibull) distribution, with burstiness more strongly predicted by semantic Class than frequency and scaling robust across languages, genres, and communication modes.
Takeaways & Limitations
Recurrence patterns link word meanings to discourse-context permutability, providing a long-term model of word usage that extends beyond Zipf’s law.
Takeaways & Limitations
The renewal model does not reproduce positive correlations between successive recurrence intervals, which are small but decay slowly with lag.
Abstract
from arXiv · showhide
Background: Zipf's discovery that word frequency distributions obey a power law established parallels between biological and physical processes, and language, laying the groundwork for a complex systems perspective on human communication. More recent research has also identified scaling regularities in the dynamics underlying the successive occurrences of events, suggesting the possibility of similar findings for language as well. Methodology/Principal Findings: By considering frequent words in USENET discussion groups and in disparate databases where the language has different levels of formality, here we show that the distributions of distances between successive occurrences of the same word display bursty deviations from a Poisson process and are well characterized by a stretched exponential (Weibull) scaling. The extent of this deviation depends strongly on semantic type -- a measure of the logicality of each word -- and less strongly on frequency. We develop a generative model of this behavior that fully determines the dynamics of word usage. Conclusions/Significance: Recurrence patterns of words are well described by a stretched exponential distribution of recurrence times, an empirical scaling that cannot be anticipated from Zipf's law. Because the use of words provides a uniquely precise and powerful lens on human thought and activity, our findings also have implications for other overt manifestations of collective human dynamics.
INTRODUCTION
Prior work identified bursty temporal patterns in natural and social events, motivating a search for analogous scaling in language. Using large text databases, the paper investigates word recurrence over timescales beyond local syntax and relates burstiness to semantics.
- Motivation: Bursty deviations from random and regular event timing suggest a dynamic counterpart to scaling laws in magnitude and frequency distributions.Language is presented as a promising domain because it is both a social activity and a medium for representing natural and biological reality.
- Motivation: Statistical natural language processing and psycholinguistics already study language dynamically, including word predictability and document-level word statistics.The non-uniform distribution of content words through texts motivates systematic study of word recurrence times.
- Data and approach: USENET discussion groups provide records of spontaneous collective language for investigating statistical questions about word usage at large scale.The study begins with discussion groups available through Google and focuses on spontaneous interactions in large communities over time.
- Contributions: Long-time word recurrence patterns follow a stretched exponential distribution because word usage contains bursts and lulls beyond syntactic timescales.The reported burstiness is driven by word semantics, while logicality or permutability makes words typically less bursty than other human activities.
- Contributions: The analysis tests generality across USENET, books of different genres, and political debates with differing levels of formality.The model addresses long timescales where local predictability and coherence studies leave off.
RESULTS AND DISCUSSION
Word recurrence times depart from Poisson expectations and are well described by a stretched exponential whose burstiness varies primarily with semantic class. A memory-based generative model links this pattern to context-dependent word use and remains broadly applicable across words and datasets, with renewal-model limitations.
- Recurrence distributions: The stretched exponential fits recurrence-time distributions better than the Poisson exponential for words such as theory and also.Observed distributions have excesses at both short and long distances and deficits near the mean recurrence time.
- Generative model: The model assumes stationary word generation in which usage probability depends on the distance since the word’s previous occurrence.Empirical power-law decay of this probability produces the stretched exponential recurrence distribution.
- Model limitation: The renewal model does not reproduce positive correlations between successive recurrence intervals, which are usually below 20% at lag one but decay slowly.These correlations mark how closely the renewal approximation matches the actual generative process.
- Semantic class and frequency: 2,128 frequent USENET words show β values from 0.2 to 0.9, with most recurrence distributions well described by the same stretched-exponential model.Words across semantic classes share the model over a wide range of scales.
- Semantic class and frequency: 103 of 116 frequency-matched verb–noun pairs have higher β for verbs, and 37 of 47 adjective–adverb pairs show higher β for adverbs.The sign tests report P ≤8 10^-19 and P ≤5 10^-5, respectively.
- Semantic class and frequency: Semantic class accounts for 0.32 of β variance, compared with 0.26 for log-frequency, while class ordering persists across inverse frequencies.Class-based grouping yields narrower within-class β distributions and better discrimination than frequency grouping.
- Interpretation: The model interprets higher semantic class as greater permutability, producing more homogeneous discourse distributions and higher β values.This connects lower burstiness to words whose meanings remain applicable across more contextual alternatives.
SUPPORTING INFORMATION
Table S1 provides detailed statistical-analysis information for all words studied across six databases.
- Table S1 contains detailed statistical information on every studied word across six databases.