Source-linked AI summary

The dynamics of correlated novelties

F. Tria, V. Loreto, V. D. P. Servedio, S. H. Strogatz

arXiv:1310.1953v1physics.soc-phcs.SI

TL;DR

The paper asks whether novelties pave the way for one another rather than appearing by chance, and models this expansion with a generalized urn. The model explains Heaps’ and Zipf’s laws through one mechanism, while the datasets display predicted correlations among novelties.

  • Problem

    The paper examines whether everyday novelties arise by chance alone or whether one novelty paves the way for another.

  • Method

    The model generalizes an urn process by conditioning on a novelty to model expansion, with reservoir size represented by U(t).

  • Results

    All the data sets display the predicted correlations among novelties.

  • Takeaways & Limitations

    The model explains Heaps’ and Zipf’s laws through the same basic microscopic mechanism.

  • Takeaways & Limitations

    The Simon model cannot reproduce empirically found values of α larger than 1, and explaining Heaps’ and Zipf’s laws is not sufficient to account for the full result.

Abstract

from arXiv · show

One new thing often leads to another. Such correlated novelties are a familiar part of daily life. They are also thought to be fundamental to the evolution of biological systems, human society, and technology. By opening new possibilities, one novelty can pave the way for others in a process that Kauffman has called "expanding the adjacent possible". The dynamics of correlated novelties, however, have yet to be quantified empirically or modeled mathematically. Here we propose a simple mathematical model that mimics the process of exploring a physical, biological or conceptual space that enlarges whenever a novelty occurs. The model, a generalization of Polya's urn, predicts statistical laws for the rate at which novelties happen (analogous to Heaps' law) and for the probability distribution on the space explored (analogous to Zipf's law), as well as signatures of the hypothesized process by which one novelty sets the stage for another. We test these predictions on four data sets of human activity: the edit events of Wikipedia pages, the emergence of tags in annotation systems, the sequence of words in texts, and listening to new songs in online music catalogues. By quantifying the dynamics of correlated novelties, our results provide a starting point for a deeper understanding of the ever-expanding adjacent possible and its role in biological, linguistic, cultural, and technological evolution.

1 Institute for Scientific Interchange (ISI), Via Alassio 11C, 10126 Torino, Italy

The paper models correlated novelties as an expanding exploration process and tests its predictions on four human-activity datasets. The model reproduces Heaps’ and Zipf’s laws while capturing temporal and semantic correlations among novelties.

  • Empirical setting: The study analyzes words, songs, Wikipedia pages, and tags, defining novelty as a first occurrence or first user-specific interaction.These four temporally ordered datasets operationalize novelty across texts, online music, Wikipedia, and social annotation systems.
  • Empirical regularities: All four datasets show sublinear growth of distinct elements, indicating Heaps’ law and a declining novelty rate over time.The growth follows D(N) ∼ N^β with β < 1, and the novelty rate decreases as t^(β−1).
  • Empirical regularities: All datasets also exhibit approximate Zipf frequency-rank tails, with exponents compatible with Heaps’ law through β = 1/α.The relation links the frequency distribution to the growth of distinct elements within each dataset.
  • Correlated novelties: Semantic novelty correlations appear through lower entropy than reshuffled data, temporal clustering, and higher within-user clustering than reshuffled sequences.The extended model reproduces the observed Heaps’ and Zipf’s laws as well as the measured behavior of S and f(l).
  • Model: The generalized urn model allows novelties to trigger further novelties by expanding the reservoir of possible elements.Conditioning the urn on a novelty is the mechanism used to model expansion of the adjacent possible.
  • Model predictions: The balance between reinforcement and triggering determines whether distinct-element growth is sublinear or linear.Sublinear growth occurs when reinforcement is stronger than triggering, whereas linear growth occurs when triggering outweighs reinforcement.
  • Implications: The same microscopic mechanism explains both Heaps’ and Zipf’s laws without requiring one law to be hypothesized in order to derive the other.The authors connect this result to quantitative testing of the expanding adjacent possible in techno-social systems.

Figures

Figures 1–3 compare empirical scaling patterns with urn-model behavior and illustrate the model’s reinforcement and adjacent-possible steps.

  • Figure 1: Figure 1 presents Heaps’ law for real datasets and the urn model with triggering, alongside corresponding Zipf’s-law plots.The real datasets include Gutenberg, Last.fm, Wikipedia, and del.icio.us.
  • Figure 2: Figure 2 displays entropy versus event count for Wikipedia, Last.fm, and the urn model with semantic triggering.The model curves average 10 realizations with ρ = 8, ν = 10, η = 0.3, and N = 10^7.
  • Figure 3: Figure 3 illustrates reinforcement and adjacent-possible steps in both the simple and semantic-triggering urn models.Semantic triggering additionally assigns labels that are conserved during reinforcement and newly introduced for adjacent-possible additions.

1 Urn model with triggering

The urn model with triggering generates sequences in which novel elements enlarge the reservoir, while repeated elements are reinforced. Both model versions predict sublinear Heaps’ growth and fat-tailed frequency-rank distributions, with analytical predictions confirmed numerically.

  • Model definition: At each step, the model randomly extracts an element, records it in the sequence, and returns it with ρ additional copies.A first-time element also adds ν + 1 brand-new distinct elements to the reservoir.
  • Model definition: The second model variant applies reinforcement only when an element is not new, changing the reservoir dynamics for first appearances.Its rule replaces the original reinforcement step with a condition that excludes first-time elements.
  • Analytical predictions: Both model versions predict Heaps’ law for the number of distinct elements and a fat-tailed frequency-rank distribution.The analysis derives formulas for the Heaps’ exponent and the asymptotic power-law exponent of the frequency-rank distribution.
  • Heaps–Zipf relation: The model’s Heaps–Zipf relation is asymptotic rather than automatic, because random sampling from a fixed power-law distribution is not assumed.The relation begins later when the Heaps’ exponent β is smaller.
  • Model validation: Numerical simulations confirm the analytical predictions for both the distinct-element growth and the frequency-rank distribution.The relation β = 1/α between Heaps’ and Zipf’s exponents holds only asymptotically, with α measured on the frequency-rank tail.

2 Detecting triggering events

The paper detects semantic triggering by measuring how occurrences of related elements cluster in sequences. Entropy and inter-event intervals are compared with globally and locally reshuffled sequences to distinguish semantic from statistical correlations.

  • Detecting semantic groups: Semantic labels group related elements, such as songs by artist or Wikipedia edits by mother page, to test whether occurrences cluster.Words are treated as their own classes because no satisfactory semantic classification is available.
  • Entropy measure: For each label A, the entropy S_A(k) measures how its k occurrences are distributed across k equal intervals after its first appearance.Uniform distribution gives maximum entropy log k; concentration in the first interval gives minimum entropy.
  • Interval measure: Triggering intervals are measured as the distribution of times between successive appearances of each label and then aggregated across labels.Shorter intervals indicate more temporally clustered repetitions.
  • Correlation controls: Global reshuffling removes semantic correlations while preserving nonstationary statistical correlations responsible for Heaps’ and Zipf’s laws.Local reshuffling from each label’s first appearance onward removes correlations between element appearances altogether.

3 The random walk model for the dynamics of novelties

The random-walk model maps semantic triggering onto exploration of an evolving graph whose topology encodes relations among elements. Its statistics are qualitatively equivalent to the semantic urn model even when graph links are probabilistic, while allowing richer relation structures.

  • Graph construction: The random-walk dynamics explore an evolving graph initialized with labeled cliques and inter-clique links formed with probability η.The walker starts randomly, then moves to neighbors or remains in place according to weight-dependent probabilities.
  • Graph construction: When a new node is visited, the model adds a clique of ν + 1 new nodes sharing a new label and connects them to existing nodes with probability η.This graph-growth rule represents an expanding adjacent possible.
  • Relation to urn model: For η = 1, the graph model maps one-to-one to the urn model with triggering; for η < 1, the correspondence is not one-to-one because graph links are fixed.The urn model instead treats transitions between elements probabilistically at each step.
  • Model behavior: Despite this difference, the two models have qualitatively equivalent statistical properties for η < 1, including Heaps’ and Zipf’s laws.The random-walk model also exhibits entropy and triggering-interval signatures.
  • Model extension: The random-walk formulation naturally supports more complex semantic relations because those relations are encoded in the growing graph topology.Different linking rules can represent richer and more realistic semantic structures.

4 Details of the datasets used

The study analyzes temporally ordered human-activity data from texts, social annotation, music listening, and Wikipedia editing. The datasets range from large aggregated corpora to selected individual users and editors.

  • Gutenberg: The Gutenberg corpus contains about 2.8 × 10^8 words from approximately 4,600 English ebooks and about 5.5 × 10^5 distinct words.The books include diverse subjects, prose, and poetry; capitalization is ignored.
  • Delicious: The Delicious dataset covers about 5 × 10^6 posts, 650,000 users, 1.9 × 10^6 resources, and 2.5 × 10^6 tags across almost three years.Post timestamps establish temporal ordering, and capitalization differences are collapsed during tag comparison.
  • Wikipedia: Wikipedia edits are represented by page, user, edit, timestamp, and mother-page identifiers, then duplicate same-user edits are removed and events are time-sorted.The non-aggregated analysis focuses on seven randomly chosen editors and screens for human contributors because active editors may be robots.

5 Results for non aggregated data

Across selected individual records, novelty growth and frequency-rank patterns broadly reproduce the aggregate findings, while triggering signatures appear at the individual level. The adjacent-possible effect is reported as stronger in collective processes.

  • Heaps’ law: Selected Gutenberg texts, Last.fm listeners, Wikipedia editors, and Delicious users show asymptotic sublinear Heaps’ growth, except that Wikipedia users grow linearly.Last.fm and Delicious curves are less smooth because users can import blocks of tracks or bookmarks, creating temporal discontinuities.
  • Heaps’ and Zipf’s relation: The model accounts for Wikipedia’s possible linear dictionary growth and predicts a connection between the Zipf exponent and the slope of that growth.This prediction is stated alongside the observed β ≃ 1 behavior for Wikipedia editors.
  • Zipf’s law: Frequency-rank distributions for words, lyrics, wiki-articles, and tags follow approximate power laws consistent with Zipf’s law.For selected texts, the measured Heaps’ exponent is not necessarily the reciprocal of the measured Zipf exponent because finite samples may not reach the asymptotic regime.
  • Triggering signatures: Entropy and interval comparisons confirm at the user level the same triggering signatures observed in the full datasets.The comparisons use original, globally shuffled, and locally shuffled sequences.
  • Individual and collective effects: The adjacent-possible mechanism operates at the individual level, and its effect is enhanced in collective processes.In Last.fm, short clusters can reflect listeners browsing several songs from the same album, lowering the associated entropy.
Loading 1310.1953v1…