Source-linked AI summary

The Oligarch Barely Steers Model Collapse in Multi-Model Ecosystems

Yangze Liu, Zhongyi Han

arXiv:2609.11146v1cs.AIcs.CLcs.LG

TL;DR

Recursive training on AI-generated text can cause model collapse, but prior multi-model studies largely assume equal contributions despite an oligopolistic market. This paper tests unequal shares in controlled five-generation ecosystems and finds that concentration barely changes collapse’s speed or destination, while the pool’s composition determines the pace.

  • Problem

    Prior multi-model research generally assumes equal data contributions, leaving open whether oligopolistic concentration changes where recursive collapse goes or how fast it arrives.

  • Method

    The study retrains 3–13-model ecosystems for five generations from shared pools mixed by market share, using natural open models and an injected probe with shares up to 90%.

  • Results

    Within the tested range, concentration barely changes collapse speed or destination, while pool composition explains drift speed with R2 = 0.68 and human text slows drift without changing its course.

  • Takeaways & Limitations

    The tested ecosystems are governed less by who holds market share than by whose text fills the shared corpus and how readily those suppliers are carried along.

  • Takeaways & Limitations

    The susceptibility index leaves more than thirty percent of drift unexplained and does not reliably order all K=13 arms.

Abstract

from arXiv · show

AI-generated text is flowing back into the training corpora of the next generation of models. Recursive training on it drives model collapse, and recent work extends the setting to many models feeding one another -- but almost always with the market split evenly, while real generative AI is an oligopoly. Concentration raises two worries: fewer, more uniform sources may make collapse faster, and later models may be dragged toward the oligarch's output. We test both in controlled ecosystems: 13 open 1--4B models form natural ecosystems of 3 to 13 players, plus an injected probe that pushes the top share to 90%; each generation, every model's output is mixed into a shared pool by market share and every model is retrained on that pool from clean base weights, for five generations. Yet within the range we test, neither worry materializes; what emerges instead is an invariance. Making the split more unequal barely changes the speed of collapse. Destinations move even less: the share and identity knobs shift five-generation endpoints by only a few percent of the drift common to all arms -- the ecosystems collapse to nearly the same place. An extreme share paired with the strongest injected bias still does not guarantee steering, and the topic shifts it does produce leave only a faint trace on the ruler that measures collapse. What sets the speed is who supplies the pool and how readily those suppliers are carried along: with every share held fixed, swapping the members of a K=3 ecosystem changes five-generation drift by 2.8x; a share-weighted index of each member's susceptibility explains the speed differences across nineteen arms with R^2 = 0.68; and replacing half the pool with human text roughly halves drift without changing its course. Within the tested range, concentration sets neither the destination nor the pace of collapse; the pace follows whose text fills the pool.

1 INTRODUCTION

This paper tests whether unequal market shares accelerate multi-model collapse or pull ecosystems toward a dominant provider. Across controlled probes and natural ecosystems, concentration barely changes either collapse speed or destination; pool composition instead determines the pace.

  • The study addresses a gap in prior multi-model work, which generally assumes equal contributions despite generative AI’s oligopolistic market structure.The paper frames concentration as two hypotheses: faster collapse from fewer sources and directional pull toward the dominant provider.
  • The experiments iterate shared-pool retraining for five generations, using natural ecosystems of three to thirteen open models alongside an injected oligarch probe.Each generation mixes outputs by preset share, then retrains every model from clean base weights.
  • Within the tested range, unequal shares, head identity, and player count leave five-generation endpoints nearly unchanged, while share concentration changes drift by less than a tenth.The ecosystems collapse toward nearly the same place despite different market structures.
  • A 90% oligarch share plus strong directional bias does not guarantee that the ecosystem follows the oligarch’s topic.Some oligarchs lose their own topic in the pool, and topic differences leave only a faint trace on the collapse ruler.
  • Across nineteen arms, a share-weighted susceptibility index explains drift speed with R2 = 0.68, identifying pool composition as the key predictor.With shares fixed, changing the members of a three-model ecosystem changes five-generation drift nearly threefold.

2 RELATED WORK

Prior work shows that heterogeneous models converge when they train on one another, but it largely assumes equal data contributions. This paper extends that setting by varying concentration, player count, and human-data fraction rather than treating dominance as binary.

  • Existing studies generally give every participant the same data contribution, leaving realistic oligopolistic market shares understudied.
  • Recent multi-model studies find convergence in output distributions or shared-task performance, but do not establish the meeting point or how to move it.
  • This paper sweeps shares from equal contribution to a single 90% supplier, varies player count from 3 to 13, and adds 0–50% human data.

3 SETUP

The study simulates repeated shared-corpus retraining in ecosystems of open language models, varying market structure, membership, player count, and human-data fraction. Collapse is measured from frozen-encoder distances between generation-specific model-output pool centroids.

  • Each generation, models generate text, outputs are mixed by preset market shares, and every model is retrained on the shared pool from clean base weights.Only text crosses generations; model weights are replaced after each five-generation cycle.
  • The injected probe uses three opposed stylistic biases to test whether an oligarch can steer the ecosystem, while natural simulations use off-the-shelf open models without injection.
  • The experiments vary share structure, head identity, player count from 3 to 13, and human-data fractions up to 50%.Natural ecosystems compare equal shares with a 28% head; the probe pushes the top share from 45% to 90%.
  • Collapse is measured as 1−cosine distance between frozen DeBERTa embeddings of model-generated pool-text centroids, excluding human lines from human-data-arm centroids.Reported distances are multiplied by one thousand.
  • Natural ecosystems use 13 open 1–4B-parameter models, three seeds, 800 texts per model per generation, and a fixed mixed pool of 2,100 texts.

4 RESULTS

Across the tested ecosystems, concentration barely changes either the destination or speed of collapse, while pool composition determines how quickly models drift. Human text slows movement along the same course, and member susceptibility predicts speed differences.

  • 4.1 THE INJECTED PROBE: THE PULL AT FULL THROTTLE DOES NOT GUARANTEE STEERING: Even a 90% share paired with strong injected bias does not guarantee that the ecosystem follows the oligarch’s topic.The strongest science bias failed to hold its topic, while an intermediate subjective bias did; two of three 90% oligarchs saw their own topics decline.
  • 4.1 THE INJECTED PROBE: THE PULL AT FULL THROTTLE DOES NOT GUARANTEE STEERING: Topic differences from the injected probe remain small on the collapse ruler: the most opposed arms end 27 apart after five generations while drifting 92 and 131.The topic-level pull is visible in the topic readout but contributes only a faint trace to geometric collapse.
  • 4.3 HUMAN DATA SLOWS COLLAPSE BUT DOES NOT CHANGE ITS COURSE: Replacing 25% or 50% of the pool with human text reduces five-generation drift from 99 to 72 or 54 without changing its course.The human-data arms follow the same direction as the oligarch arm, corresponding roughly to one or two fewer generations of travel.
  • 4.4 HOW FAST AN ECOSYSTEM COLLAPSES DEPENDS ON WHOSE TEXT FILLS THE POOL: A share-weighted index of member susceptibility explains five-generation drift across nineteen arms with R^2 = 0.68.The index measures how readily, on average, the text in the pool carries a model along.

5 DISCUSSION AND CONCLUSION

Within the tested range, concentration changes neither the destination nor the pace of recursive collapse; speed instead follows whose text fills the shared pool. The study’s conclusions remain bounded by model scale, player count, encoder dependence, five-generation observation, and a post-hoc speed fit.

  • Discussion and conclusion: The share and identity manipulations leave endpoint differences at only a few percent of the common five-generation drift, while the shared collapse remains much larger.The paper reports a share-induced generation-5 displacement of 0.8 versus five-generation movements of 92 and 99.
  • Discussion and conclusion: Pool composition, rather than concentration itself, predicts collapse speed; a share-weighted susceptibility index is the only tested single-number predictor.Human text and text from hard-to-carry models alter composition and slow the walk, while the destination remains nearly fixed.
  • Discussion and conclusion: The main scope boundary is that experiments cover five generations, 1–4B models, and at most thirteen players, while larger-scale behavior remains unresolved.A one-time 7–8B probe could not distinguish under-convergence from a scale effect.
  • Discussion and conclusion: The composition–speed relationship explains substantial but incomplete variation, leaving more than thirty percent of variance unexplained and clustered by ecosystem size.The paper therefore presents it as a coarse relationship rather than a law of speed.
  • Discussion and conclusion: The findings use public open-weight models and corpora, exclude human subjects and personal data, and make no policy recommendation.Code, configurations, and per-arm measurements are slated for release.

A METHODS AND REPRODUCIBILITY DETAILS

The experiments use nested ecosystems of 13 open 1–4B models, controlled market shares, repeated shared-pool retraining, and frozen-encoder measurements. Reproducibility details specify the models, sampling, training, generation, measurement, and equivalence analysis.

  • Models and ecosystem construction: Thirteen open 1–4B models form nested K=3, 5, 8, and 13 ecosystems, with Phi-2 designated as the 28% head because its initial output is most distinguishable.Identity-counterfactual arms instead assign the head share to SmolLM2 or Qwen3.
  • Experimental design: Natural arms compare uniform shares with a 28% head, while identity, player count, human-data fraction, and injected-probe concentration provide the main experimental variations.Natural ecosystems stop at 28%; the 90% extreme is confined to the injected probe.
  • Per-generation pipeline: Each generation produces 800 texts per model, samples a fixed 2,100-text pool by share, fully fine-tunes every model from clean base weights, and repeats for five generations.Three seeds share generation-0 outputs and prompts within paired comparisons.
  • Measurement: The main geometric measurement embeds generated text with a frozen DeBERTa encoder, computes centroid distances as 1−cos per seed, and averages across seeds.Human training lines are excluded from centroid measurement; reported distances are multiplied by one thousand.
  • Equivalence analysis: Equivalence bounds for uniform versus 28% arms are post-hoc, based on three seeds, and express endpoint separation as a fraction of paired five-generation drift.The reported upper-bound fractions are 7.2%, 2.3%, and 1.2% for K=5, 8, and 13.
  • Measurement: The encoder-free degradation analysis contains 330 trajectories and reports per-generation medians with interquartile ranges for perplexity, Distinct-4-gram fraction, and Self-BLEU-4.Perplexity uses a frozen gpt2-large and a log axis; each metric reads each model’s newly generated text.

B INJECTED-PROBE SUPPLEMENT

The injected probe finds that even extreme oligarch shares and strong topic biases do not reliably steer the ecosystem or accelerate collapse. Topic persistence varies by bias, while geometric collapse remains dominated by common drift.

  • Injected bias: The generation-5 topic outcomes are non-monotone: the strongest science bias fails to hold its topic, the intermediate subjective bias holds it, and the weakest factual bias shifts toward subjectivity.The factual topic reaches a 0.71 share under the subjective direction at generation five.
  • Results: 194 versus 89–131: the 90% factual arm has the largest five-generation drift, while 90% science drifts 92 and the uniform arm 112.Across the reported arms, the strongest geometric collapse is associated with the weakest factual injection, not a simple concentration effect.
  • Geometric separation: Generation-5 between-arm distances span only 5–28, below the smallest within-arm drift of 89, so topic differences leave a faint geometric trace.The most topic-opposed pair is separated by 27 on the between-arm ruler.
  • Concentration: The 45%–90% science-share sweep ranges from 89 to 98, versus 112 for the uniform arm, showing no acceleration with concentration.With the head identity fixed, increasing its share mildly decelerates rather than accelerates drift.

C HUMAN-DATA DOSE RESPONSE AT K=3

Adding human text to the shared pool consistently raises the two model-output diversity readings at K=3. Distinct-2 increases monotonically with dose, whereas dispersion shows a non-monotone response, and a third cross-model measure is too noisy to resolve ordering.

  • Dose response: +0.013/+0.029/+0.054: distinct-2 rises monotonically with 5%, 10%, and 25% human-text doses.All three seeds agree in sign at every dose.
  • Dose response: +0.0072/+0.0061/+0.0105: dispersion increases across doses but is not monotone, with 5% and 10% indistinguishable at this noise level.The paired differences are positive for all doses and seeds, but the dose means do not form an increasing sequence.
  • Interpretation: Human text consistently lifts both readings across all doses and seeds, matching the slowdown direction reported for the larger K=13 experiment.The measurements use each model’s own generated text rather than the diluted mixed pool.
  • Limitation: 6.4× seed spread makes the cross-model distance unreadable: its uniform-arm values range from 0.0136 to 0.0873, while the largest paired difference is 0.078.The point estimates lean negative at 10% and 25% in two of three seeds, so this measure cannot establish the intended ordering.

D COMPOSITION INDEX DETAILS

The composition index explains collapse speed better than concentration, player count, or initial similarity, but its fit has systematic exceptions and important measurement caveats. Membership can change drift substantially even when shares and HHI are fixed.

  • Measurement caveats: B reaches R^2 = 0.773 without the frozen encoder, but this is not independent validation because drift is still measured by that encoder.C is near-tautological, while D becomes contentful only after dividing out starting distance.
  • Within-tier checks: Within tiers, the index orders all K=3 pairs, but shows no ordering among K=13 arms; the within-tier claim therefore holds only for K=5 and K=8.The reported rank correlations are +1.000 for K=3, +0.800 for K=5, +1.000 for K=8, and +0.200 for K=13.
  • Caveat: The index misses a systematic K=8 elevation: all five K=8 residuals are positive, and one K=8 indicator raises R^2 from 0.680 to 0.766.The authors do not identify what the index omits.
  • Composition index: R^2 = 0.680 and rank correlation 0.811: the share-weighted travel propensity predicts five-generation drift across nineteen arms.Out-of-sample performance remains substantial but weaker, with Q^2 = 0.552 and rank correlation 0.791.
  • Comparators: Under 0.01 of drift variance: numeric player count and HHI explain almost nothing, versus 0.680 for composition.Categorical player-count and HHI fits reach only adjusted R^2 values of 0.06 and below zero.

E A TOY MODEL OF SHARES VERSUS THE FADE

A toy dynamical model explains why share can strongly affect imitation speed yet barely move the endpoint when common fade dominates. The model predicts that endpoint displacement from shares is small relative to the shared collapse trajectory.

  • Dynamics: The update averages the shared pool, each model’s factory position, and a common fade direction, with weights summing to one and no overshoot.The pool is the share-weighted average of model positions, while clean-base retraining reintroduces the factory prior.
  • Toy-model result: At fixed fade fraction λ/(1 −α), imitation strength α changes how fast positions settle, not where they end.The remaining outcome is determined by the fade fraction λ/(1 −α), not by imitation strength alone.
  • Zero-fade limit: Under the toy’s zero-fade limit, the ecosystem center of mass settles at the share-weighted average of factory positions.This is the simple mechanism behind the policy concern that imitation could favor the largest contributor.
  • Empirical scale: 0.6 versus 99: changing one model’s share to 28% shifts the initial pool centroid by only 0.6, while five-generation collapse moves 99.The share-induced displacement is under 1% of the collapse displacement on this ruler.

F ENCODER CHECKS

Audit encoders qualitatively replicate the endpoint, between-arm separation, and human-data slowdown results, but cross-domain angle readings are encoder-specific and not load-bearing.

  • All four encoders agree qualitatively that collapse endpoints lie farther from the human-corpus centroid than generation-0 outputs.
  • 0.19–0.56 of the seed-noise floor: between-arm separation stays below that floor across all 12 encoder-by-K comparisons.Energy-distance ratios likewise remain below one at 0.23–0.59 across all 12 comparisons.
  • 1.7–2.7× slower: replacing recursive-only training with 50% human data reduces drift across all three audit encoders.The dose ordering also holds consistently: 50% is slower than 25%, which is slower than none.
  • Cross-domain angle readings flip sign on all three audit encoders, so that directional claim is specific to DeBERTa.The paper retains distance and speed measures as its load-bearing cross-encoder checks.

G THE 7–8B PROBE

A larger-model probe shows trajectories continuing to close, but its setup cannot distinguish a scale effect from under-convergence.

  • 0.99 to 0.40: the three-arm separation-to-drift ratio falls through generation 10, indicating that trajectories continue closing.
  • The 7–8B probe is exploratory rather than a scale test because it uses one seed, a different generation backend, and an under-converged training budget.These factors make slower convergence and a scale effect indistinguishable.

H GENERATED-TEXT SAMPLES

Generated-text samples show that later outputs retain much of their syntax while losing factual reference and becoming increasingly repetitive or incoherent across models.

  • The sample contains 27 continuations selected from nine prompts across three models and three generations, with none omitted.The selection rule was fixed in advance to prevent post-hoc cherry-picking.
  • The transcription applies five declared normalization and masking operations, including 300-character truncation and whitespace normalization.
  • By generation 5, syntax mostly survives while word-to-word constraints dissolve into fixed lists, synonym chains, or character noise.Qwen2.5 repeats a prompt-independent word skeleton, granite produces synonym chains or noise, and SmolLM2 loses reference despite retaining sentence shape.
  • Generation-0 continuations resemble encyclopedia entries with wrong facts, while syntax, register, and topic remain broadly intact.One example incorrectly assigns Pharomachrus, the quetzal genus, to the Australian ringneck.
Loading 2609.11146v1…