Source-linked AI summary

Will we run out of data? Limits of LLM scaling based on human-generated data

Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, Marius Hobbhahn

arXiv:2211.04325v2cs.LGcs.AIcs.CLcs.CVcs.CY

TL;DR

The paper asks whether limited public human-generated text could constrain further LLM scaling. It projects training-data demand against the available stock, finding likely exhaustion between 2026 and 2032, while examining data-efficiency improvements, transfer learning, and synthetic data as possible ways forward.

  • Problem

    The paper asks whether the limited availability of public human text data could constrain further LLM scaling.

  • Method

    The paper projects growth in LLM training-dataset sizes and the stock of available public human-generated text, while examining possible ways to circumvent bottlenecks.

  • Results

    Models may use the full supply of public human text between 2026 and 2032, or one or two years earlier if frontier models are overtrained.

  • Takeaways & Limitations

    The current paradigm based on public human text may not continue a decade from now, but alternative data sources may allow ML systems to keep scaling.

  • Takeaways & Limitations

    Synthetic-data effectiveness is mixed because repeated training can produce homogeneous outputs, diminishing or negative returns, and worse scaling behavior.

Abstract

from arXiv · show

We investigate the potential constraints on LLM scaling posed by the availability of public human-generated text data. We forecast the growing demand for training data based on current trends and estimate the total stock of public human text data. Our findings indicate that if current LLM development trends continue, models will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained. We explore how progress in language modeling can continue when human-generated text datasets cannot be scaled any further. We argue that synthetic data generation, transfer learning from data-rich domains, and data efficiency improvements might support further progress.

1. Introduction

Recent LLM progress depends on increasingly large human-generated text datasets, while the available public supply is finite. The paper examines when this growing demand may exhaust that supply and what alternatives could support continued scaling.

  • The largest public human-text datasets already contain tens of trillions of words from billions of web pages.
  • Increasing training-dataset size is crucial for efficiently improving LLM performance under neural scaling laws.
  • The paper models training-data demand and public human-text production to predict when LLM development could exhaust the available stock.
  • The paper explores synthetic data, transfer learning from data-rich domains, and non-public data as possible ways to circumvent the constraint.
  • The analysis responds to concerns that insufficient data availability could limit machine-learning progress.

2. A model of data scarcity

The paper models both the available stock of public human text and the dataset sizes demanded by LLM training, incorporating tokenization, data quality, repetition, and compute constraints. Its projections indicate that public text could become a major bottleneck this decade, with exhaustion centered on 2028 and occurring earlier under intensive overtraining.

  • Model variables: The model distinguishes the total public human text stock from the quantity used in an LLM training dataset.Dataset size is defined as the number of tokens in the training dataset of interest.
  • Data stock: 3100T tokens is the projected total stock of internet text, including indexed and deep-web data, with a 95% confidence interval of 1900T–5200T.Future accumulation is projected by scaling annual 2024 uploads with the number of internet users and cumulatively summing contributions.
  • Responses to scarcity: Compute constraints could slow historical dataset-size growth, while synthetic data, transfer learning, non-public data, and data-efficiency techniques are proposed as ways to continue progress.The paper identifies energy efficiency, data-center electricity, chip capacity, and economic constraints as limits on compute scaling.
  • Dataset demand: 2028 is the median projected exhaustion year for public text data, and exhaustion becomes very likely by 2032 if past trends continue.At exhaustion, models are projected to use around 5e28 FLOP during training.
  • Dataset demand: 5x overtraining would create the data bottleneck one year earlier, at approximately 6e27 FLOP.Overtraining increases effective data use through repeated training epochs, while the paper notes that its quality adjustment relies on a nonstandard measure of data quality.

3. Beyond public human text data

Because public human text may become a bottleneck, the paper evaluates synthetic data, transfer from other domains, non-public data, and data-efficiency techniques as possible alternatives.

  • Three broad strategies are highlighted: model-generated data, multimodality and transfer learning, and non-public data.
  • AI-generated data: Synthetic data could expand training data dramatically, but repeated use may produce homogeneous outputs, diminishing returns, or worse scaling.Greater diversity and mixtures of human-generated and synthetic data may mitigate these problems.
  • AI-generated data: Synthetic data has shown promise in mathematics, programming, and games, where outputs are relatively easy to verify.Examples include AlphaZero self-play and AlphaGeometry’s synthetic geometry-problem data.
  • Transfer learning and multimodality: Image and video stocks are not large enough to prevent a bottleneck, but data-rich domains such as financial and scientific datasets may provide substantially more data.Some multimodal data shows synergy with text, although leveraging other domains is not always clear.
  • Non-public data: Non-public platforms and messaging applications may contain about 3 quadrillion tokens, delaying a bottleneck by roughly one and a half years, but privacy, quality, and fragmentation limit usefulness.
  • Data efficiency: Data-efficiency improvements may compensate for exhausted data stocks, although the fraction of efficiency gains attributable to using less data remains uncertain.Reported LLM efficiency improvements are 0.4 OOM/y [95%: 0.1, 0.8].

4. Discussion

The paper concludes that scaling based on public human text is unlikely to continue for another decade, while alternative data sources may be adopted before then. It also notes that quantitative understanding of transfer learning and synthetic data remains incomplete.

  • Transfer learning and self-generated data are identified as promising pathways beyond public human text constraints.
  • The current paradigm based on public human text is expected to become unsustainable within a decade, while alternative data sources may allow scaling to continue.
  • Quantifying transfer-learning and synthetic-data benefits requires better understanding of data quality, distribution proximity, and cross-distribution synergy.
  • The analysis does not address how desired capabilities determine data needs or how autonomous real-world exploration could change future data sources.

5. Conclusion

The analysis projects that public human text data could become exhausted for LLM training between 2026 and 2032, though data-efficiency improvements and alternative data sources may help overcome the bottleneck.

  • 2026–2032: models may use the full supply of public human text data if rapid dataset growth continues.The exhaustion point may arrive one or two years earlier if frontier models are overtrained.
  • Data-efficiency improvements, transfer learning, and synthetic data generation may help overcome the public human text bottleneck.
  • Long-term projections are uncertain because AI advances rapidly and key parameters, including data-efficiency growth and emerging-method gains, require further study.

Impact Statement

The paper highlights justice, compensation, privacy, and security issues arising from large-scale use of human-generated and platform data for AI training.

  • Compensating creators of scraped training data is presented as an important fairness and justice consideration.
  • Social media and messaging-app data could be valuable for AI training but raise serious privacy and security concerns.Sensitive personal information could be exposed without proper safeguards.

A. Theoretical growth model of the web

The web-growth model combines population growth with internet-penetration growth to estimate how human-generated data production changes over time. An exponential-times-sigmoid model better captures observed deceleration in Reddit submissions.

  • The exponential-times-sigmoid model captures deceleration in Reddit submission growth better than purely exponential or purely sigmoidal alternatives.The exponential model misses decreasing growth rates, while the sigmoidal model plateaus at zero growth.
  • Reddit submission data are too short to make the model’s additional deceleration from subexponential population growth noticeable.
  • The model uses population projections and a sigmoid internet-penetration function to estimate data production over time.Internet penetration is modeled from approximately 0% in 1950 to 50% in 2016, with 0.15 as a fitted scale parameter.

B. Estimating the size of the indexed web

The paper estimates the indexed web using search-engine results and Common Crawl statistics, while examining uncertainty from temporal variation, language distribution, and possible index bias.

  • Pivot-word frequencies from RefinedWeb are combined with Google’s reported result counts, then averaged across words for a more robust index-size estimate.
  • Figure 7 compares Reddit submission-growth functions on linear and logarithmic scales; the linear plot favors the sigmoid-times-exponential model for recent years.
  • 330B web pages: the estimated mean Google index size from 100 pivot words, with a 250B median and a 100B–1200B 95% CI.The estimates are approximately log-normal and are about four times larger than Common Crawl’s 75B unique URLs.
  • Google-index estimates are substantially higher than earlier estimates because the study uses the Google Custom Search JSON API rather than web-interface counts.The web-interface numbers are around half as large for the same search terms.
  • A change in web word frequencies across 2013–2021 altered the resulting estimates by less than 10%.The paper therefore considers this source of temporal noise unlikely to be significant.
  • The exact language distribution of the web is uncertain, and Google and Common Crawl may underrepresent non-Western content.A lower English share could increase the total-size estimate, but the paper states that this would not significantly change its conclusions.
  • The indexed web’s stable temporal mean despite growth in users and Wikipedia motivates hypotheses involving link rot, index-size constraints, and systematic estimation bias.

C. Non-public text data

The paper estimates non-public text stocks from major closed platforms and finds that deep-web text is roughly comparable to indexed-web text. Instant messaging and email add substantial tokens, but their practical effect on the bottleneck is limited by quality and duplication concerns.

  • Facebook produces about 10T tokens annually, yielding an estimated stock of around 100T tokens over roughly ten years.
  • Instagram generates around 800B tokens annually, corresponding to approximately 8T tokens over ten years.
  • Twitter generates about 1.5T tokens per year, with an estimated total stock of around 17T tokens over twelve years.
  • Reddit contains approximately 75B tokens per year and about 600B tokens in total based on posts and comments.
  • The deep web is roughly comparable in size to the indexed web, so using it would delay a data bottleneck by only a couple of years.
  • Instant messaging adds about 50% to the stock of text at face value, delaying a data bottleneck by less than one year, while its training usefulness is poorly documented.
  • Assuming 10% of emails are unique and average 50 words, email contributes roughly 625T tokens, comparable in order of magnitude to deep-web text.

D. Non-text data

The paper evaluates images, video, and exotic modalities by converting them into text-token equivalents. Images and video appear unlikely to materially change the conclusions, while exotic modalities remain difficult to assess because of redundancy, noise, and uncertain synergy with text.

  • A decade of image production corresponds to a few hundred trillion text-token equivalents, roughly the scale of raw Common Crawl.
  • Including images or video is not expected to produce large changes because each modality’s estimated stock is similar to web data.
  • A decade of YouTube video corresponds to roughly one quadrillion text-token equivalents after scaling by YouTube’s traffic share.
  • Image and video production is easier to scale than text; existing image sensors could theoretically produce 2e17 seconds of video in one year.
  • Astronomy and genomics could represent roughly 1e18 tokens, but high redundancy, noise, and uncertain synergy with text prevent evaluating them as lasting data sources.

E.2. Tokenization Using GPT2 Tokenizer Across Various Datasets

Across datasets, GPT2 and other modern tokenizers produce broadly consistent token counts per character. This supports using approximate word-to-token conversions across studies despite tokenizer differences.

  • GPT2 produces 2.22 to 4.15 characters per token across diverse datasets.
  • Modern tokenizers remain between 2 and 5 characters per token across selected datasets, indicating little effect of tokenizer choice on tokens per character.
  • The consistency of characters per token supports treating word-to-token conversion as roughly independent of tokenizer choice.

F. Overtraining in the context of data scarcity

The paper models profit-maximizing scaling under data scarcity by combining a parametric loss scaling law with inference demand and compute constraints. It finds that optimal policies can become data-constrained, and greater overtraining exhausts the available data stock sooner.

  • The analysis examines how data scarcity affects optimal scaling decisions, including whether developers should overtrain models.
  • Chinchilla scaling uses a compute-optimal dataset-to-parameter ratio D/N of around 20.
  • Overtrained models have D/N above the Chinchilla-optimal ratio, requiring more training data but less inference compute at fixed training compute.
  • The model maximizes profit over parameters, dataset size, inference price, and inference demand under a joint training-and-inference computational budget.
  • For compute-optimal loss scaling of approximately C^-0.15, profit increasing with training compute requires demand sensitivity r to exceed about 7.
  • Higher r and h increase returns to overtraining; r = 7, h = 2 produces about 5x overtraining, while r = 10, h = 2.4 yields overtraining that grows with compute.
  • Under some assumptions, optimal scaling is data-bottlenecked rather than compute-bottlenecked, and more overtraining uses the data stock completely earlier.

G. Limits of undertraining in a data bottleneck

With a fixed data stock, increasingly large models can be undertrained to extract further gains, but only at sharply higher compute cost and for a limited period.

  • Undertraining increasingly large models on a fixed data stock can provide additional performance as compute increases.The approach adds parameters while keeping training data fixed, using a parametric scaling law to predict reducible loss.
  • Up to 2 additional orders of magnitude of compute-optimal scaling may be obtained through undertraining.This estimate assumes the training data is fixed at 300T tokens.
  • 2-3 orders of magnitude more compute are required to obtain that undertraining benefit.
  • The strategy can sustain a decreasing rate of progress for 3-6 additional years before the final plateau.The model compares this fixed-data scenario with compute-optimal scaling under unlimited data.
Loading 2211.04325v2…