Source-linked AI summary

Children, but not language models, show accelerating returns in word learning

Michael C. Frank

arXiv:2608.17120v1cs.CL

TL;DR

The paper asks whether children’s vocabulary growth accelerates with linguistic experience and tests this against language models using comparable learning-return measures. It finds accelerating returns in children but constant proportional returns in language models.

  • Problem

    The study addresses whether children’s vocabulary growth reflects accelerating rather than constant returns on linguistic experience, including relative to language models.

  • Method

    The paper models vocabulary growth with an accelerating accumulator and estimates comparable learning slopes for children and language models.

  • Results

    Children show increasing returns from additional experience, whereas language models show constant proportional returns without acceleration.

  • Takeaways & Limitations

    Acceleration in learning from experience is a fundamental difference between children and language models and warrants further investigation in artificial systems.

  • Takeaways & Limitations

    The evidence relies on statistical fit and model comparison rather than causal manipulation, so the accumulator remains untested as a mechanistic account.

Abstract

from arXiv · show

Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed. Prior models describe vocabulary growth as evidence accumulation over time. Here we show that the process is best characterized as accelerating accumulation: children learn more from each additional unit of linguistic experience than they did from the one before. In contrast to children, language models -- even those trained on child-directed speech -- do not accelerate. Instead, they show constant proportional returns on new data, consistent with scaling laws. Children learn using many orders of magnitude less training data than language models; their increasingly efficient use of their learning input is a candidate explanation.

An accelerating accumulator model of vocabulary growth

The accelerating accumulator model represents vocabulary learning as word-specific accumulation shaped by child ability and word difficulty. Unlike a pure accumulator with linear returns, it allows later language input to contribute more through an acceleration exponent κ > 1.

  • Model framework: The model treats each word as a bucket filled by input tokens, with learning occurring when its bucket is full.Accumulator models explain vocabulary spurts while accommodating word-level differences such as concreteness and syntactic category.
  • Model framework: The Rasch formulation models the probability of producing word j from child ability θ_i and word difficulty δ_j.The variables are log-transformed because words heard per hour and vocabulary size have true zeros.
  • Accelerating accumulator: The acceleration exponent κ determines how learning returns change over development: κ = 1 gives linear returns, whereas κ > 1 makes later input more valuable.The child’s ability combines baseline efficiency ξ_i, developmental accumulation over log(t/a_0), and waking hours H.
  • Model extension: The proposed model extends longitudinal growth models by allowing children to vary in both baseline accumulation efficiency ξ_i and acceleration κ_i.It is close to earlier formulations but had not previously been proposed in this form.

Fitting the accumulator model

Using item-level longitudinal CDI data from five datasets, the study fit an accelerating accumulator model and compared it with ablations. The accelerating model best fit children’s vocabulary growth, with strong acceleration across all datasets and substantial between-child variation.

  • Fitting the accumulator model: The analysis used item-level longitudinal CDI data from 1,841 monolingual, typically developing children aged 8–36 months across five datasets and fit the models with Bayesian inference.Three datasets were from American English, with one each from Norwegian and Japanese; model ablations removed individual acceleration, individual efficiency, or acceleration entirely.
  • Fitting the accumulator model: Acceleration estimates ranged from 10.6 [10.5, 10.8] to 13.3 [13.2, 13.3], robustly rejecting the pure accumulator value κ= 1 in every dataset.Estimates remained essentially unchanged with a larger dataset, hierarchical pooling, and a 2PL IRT model.
  • Fitting the accumulator model: Between-child standard deviations in acceleration ranged from 3.2 to 6.6, with 99.5% of the fitted population above the pure-accumulator value κ= 1.Sparse CDI series could not reliably predict individual acceleration beyond the population mean.
  • Fitting the accumulator model: The accelerating accumulator provided the best fit to children’s growth trajectories, improving item-level expected log predictive density over the next-best model by >760 in all five datasets.Its fits closely matched curves from a flexible non-parametric GAMLSS beta regression used for CDI percentile norms.

Accelerating accumulation implies less exposure is needed for later-acquired words

As children develop, they need dramatically fewer exposures to learn words, with exposure requirements declining exponentially across word classes. The accelerating accumulator captures this pattern, whereas a pure accumulator predicts a much shallower decline.

  • Accelerating accumulation implies less exposure is needed for later-acquired words: Exposure requirements decline exponentially with age across nouns, verbs, and function words under the accelerating accumulator, unlike the much shallower decline predicted by the pure accumulator.The pure accumulator predicts that many early-acquired words would not be learned until adolescence.
  • Accelerating accumulation implies less exposure is needed for later-acquired words: Children may hear “book” thousands of times before producing it, but later-acquired words such as “ibex” can be learned after only a handful of exposures.
  • Accelerating accumulation implies less exposure is needed for later-acquired words: The estimated efficiency increase is likely understated because exposures vary in informativeness, including whether speech is directed to the child or supported by physical context.

Language models do not accelerate

Language models show no acceleration in vocabulary growth: word-level learning accumulates gradually, unlike the accelerating pattern observed in children. Their learning trajectory is evaluated through acceleration exponents and proportional returns on new data.

  • Language models do not accelerate: Language models show no acceleration in vocabulary growth, with acceleration exponents centered near 1 and gradual accumulation of knowledge about individual words.This pattern was observed in GPT-2 models trained on 24M words of English child-directed speech, approximating a three-year-old child’s total language input.
  • Language models do not accelerate: Word-level acceleration was estimated by fitting sigmoidal curves to surprisal trajectories across training, whose slopes measure how quickly evidence for each word accumulates.The sigmoid slopes have the same units as the acceleration parameter κ in the child vocabulary-growth model.
  • Language models do not accelerate: The analysis also compared learners’ proportional returns on new data by measuring how much remaining possible loss is removed by additional training data.This comparison links language-model scaling-law analyses of loss and data entropy to the corresponding quantity computed for children.

Discussion

Children’s vocabulary learning accelerates with experience, unlike language models, which show constant proportional returns. The discussion attributes this difference to developmental and inferential changes and identifies mechanistic comparison between children and artificial systems as a priority.

  • Children need fewer exposures to acquire each new word later in development, whereas language models show constant proportional returns, even when trained on child-directed data.Converted to the same metric, children’s returns from additional experience increase while conventional language models’ returns remain constant.
  • Developmental gains in learning, memory, processing speed, and word recognition may help children extract more signal from each utterance (24, 37–39).In model terms, older children’s learning “buckets” require fewer drops to fill, producing acceleration.
  • Children may also accelerate by learning to learn: language knowledge, generalizations, and inferential abilities help them extract more information from new utterances (40, 41, 42, 43).They leverage word meanings, syntax, and pragmatic inference to identify gaps in their knowledge (3, 5).
  • Simple changes in input quantity or ordering are unlikely to explain acceleration, although richer input or child–caregiver interaction could contribute (44, 45, 36, 46).Parents generally do not fine-tune language to developmental level, and curricularizing input by age or simplicity does not improve language-model performance.
  • Language-model comparisons remain difficult because models lack developmental maturation, and analogous in-context learning typically requires vastly more data than children receive (47, 48).Few-shot rapid word learning may arise through meta-learning rather than pretraining (49), motivating linking hypotheses between children and language models (50).
  • The accelerating accumulator is supported by statistical fit and model comparison, but further work must test whether it is a mechanistic account.The study also relies on parent-report data from the validated CDI measure, and future work should characterize when artificial systems can produce accelerating returns.

List of Supplementary Materials · Supplemental Information · Materials and Methods

The supplementary methods describe longitudinal vocabulary-data filtering, Bayesian model fitting, learning-efficiency estimation, language-model training, and matched measures of proportional return and acceleration. These procedures compare children’s vocabulary learning with language models using aligned input-based metrics.

  • Vocabulary data: Children’s longitudinal vocabulary datasets included all children with at least three observations, while quality filtering excluded administrations with implausibly large decreases or increases.The filter removed 36 children, or 1.9% of the total.
  • Bayesian models: Bayesian models represented individual variation around population efficiency and acceleration, using weakly informative priors and convergence requirements for key parameters.Models were fit in Stan with four chains, 1,000 warmup iterations, 1,000 sampling iterations, and requirements of R-hat < 1.01 and bulk ESS above 400 for primarily interpreted parameters.
  • Learning efficiency: Word-level acquisition ages were computed from weighted English M3 difficulty estimates, and monthly exposures combined input rate, acquisition age, and corpus word probability.Input rates came from prior estimates, while word probabilities were computed from North American English corpora in childes-db release 2021.1.
  • Language model training: Language models used randomly initialized GPT-2-small models trained on child-directed English data, with standardized BPE tokenization and surprisal measured for 609 CDI words in held-out contexts.The primary CHILDES preparation contained 24M words, approximately 47M tokens; training and dataset-size analyses varied random seeds and used nested dataset sizes.
  • Proportional Return on Input: Language-model proportional return was defined from a power-law surprisal fit, yielding a constant γLM = β ≈ 0.32 across every input budget.The fitted scaling law had E ≈ 2.94 nats, B ≈ 312, β ≈ 0.32, and R2 = 0.998.
  • Proportional Return on Input: Children’s proportional return was defined from vocabulary-production loss and rises toward each child’s κ_i as vocabulary learning saturates.The marginal return is derived from the change in production loss with respect to log time, linking proportional return to the remaining probability of not producing words.
  • Computing Acceleration in LMs: Acceleration was compared by matching children’s κ_i to the logit slope of word learning, estimated for language models by fitting logistic curves to word surprisal across log10 training data.For language models, the corresponding slope is 1/(scale_w ln 10), expressed as logits per unit of natural-log input.

Supplementary Text … Datasets

The supplementary analyses relate the accelerating accumulator to prior models, show that acceleration cannot be identified from a single cross-sectional observation, and document dataset selection and quality-control procedures. Robustness analyses include children with two or more administrations, while the main analysis focuses on children with at least three.

  • Comparison with Prior Models: Prior accumulator models arise as special cases of the accelerating accumulator by fixing input rate, learning efficiency, or acceleration components.The model decomposes latent ability into child-specific input rate and learning efficiency, with prior models fixing one or more components.
  • Comparison with Prior Models: Models (20), (8), and (9) correspond respectively to M0, M1, and M2, differing in whether acceleration and individual learning efficiency vary.Model (20) is a pure accumulator; (8) introduces growth proportional to t^(D+1), equivalent to κ; and (9) varies α but not κ.
  • Comparison with Prior Models: Leverage can alter vocabulary-growth shape and timing but cannot generate acceleration absent acceleration in the distribution of word difficulties.A related pure accumulator with an additional per-word accumulation-speed parameter maps onto M0 combined with the fitted 2-parameter logistic variant.
  • Non-identifiability in Cross-Sectional Data: A single cross-sectional observation cannot identify a child’s acceleration because the same observation may reflect either efficiency ξ_i or acceleration κ_i.Figure 4 illustrates this non-identifiability using simulated data consistent with stable between-child intercept differences or varying growth slopes.
  • Datasets: The datasets include children with three or more administrations in the main analyses, while supplementary analyses test robustness using children with two or more administrations.Table 1 reports dataset characteristics, including analysis-sample sizes and ages under both exclusion criteria.
  • Datasets: The supplementary dataset documentation distinguishes analysis samples by exclusion criterion and reports median administrations per child.These characteristics are summarized across all datasets in Table 1.
  • Datasets: Quality checks excluded a small portion of observations, likely because of data-entry issues or faulty correspondences in archival datasets.Figure 5 marks flagged administrations and children excluded entirely in red, while retained children are shown in gray for reference.

Full Model Comparison Results · Individual-level Cross-validation

Across both datasets, Table 2 compares full models using leave-one-out predictive accuracy relative to M3, the best model. Individual-level cross-validation shows that prediction depends on the number of observations, while acceleration remains predictive of trajectories but is not itself tested by the M2–M3 comparison.

  • Full Model Comparison Results: Table 2 reports leave-one-out model comparisons across both datasets, with ELPD differences measured relative to M3, the best model.More negative ELPD differences indicate worse performance.
  • Individual-level Cross-validation: The prospective test used 371 Norwegian children with at least six longitudinal CDIs to predict each child’s next administration from their most recent k observations.The three-month-ahead design avoids confounding data density with the gap between fitted and predicted observations.
  • Individual-level Cross-validation: With k = 2, M2 clearly outperformed M3 because M3 fully parameterized two observations per child and absorbed measurement error.M2 is the fixed-slope model, whereas M3 includes a per-child slope.
  • Individual-level Cross-validation: The M2–M3 performance gap decreased as k increased, and the models were tied by k = 5.With denser data, separate child-specific slopes may yield better individual fits, but the shared-slope model made acceptable predictions in practice.
  • Individual-level Cross-validation: Because both M2 and M3 include population acceleration, this comparison does not test whether acceleration exists.Precisely estimating each child’s acceleration requires sustained, precise longitudinal measurement.
  • Individual-level Cross-validation: M2 predicted held-out administrations better than M20 at every data depth, supporting acceleration as predictive of individual trajectories.M20 constrains the acceleration exponent to κ = 1 while retaining each child’s efficiency term ξ_i.
  • Individual-level Cross-validation: The M2-versus-M20 test isolates whether a free population acceleration exponent improves individual-trajectory predictions.M2 retains the free acceleration exponent, whereas M20 removes it; positive dELPD/child favors acceleration.

Comparison to a Non-Parametric Quantile Model … Linear vs. Logarithmic Age

Across five datasets, the accelerating accumulator closely reproduces the quantile structure captured by a flexible non-parametric model. Its parameter estimates remain robust across modeling choices and data-quality filters, while logarithmic-age accumulation is preferred to linear-age accumulation.

  • Comparison to a Non-Parametric Quantile Model: The non-parametric comparison used GAMLSS beta-regression with penalized splines to model the proportion of words produced by age.This flexible family is used to produce normative curves for the CDI.
  • Comparison to a Non-Parametric Quantile Model: Across all five datasets, M3 tracks the non-parametric GAMLSS fit and empirical quantiles, providing a faithful low-dimensional compression of the data’s quantile structure.GAMLSS models the proportion of words produced using beta regression with penalized splines on the mean and scales.
  • Robustness Across Model and Data Settings: LOO fits and κ estimates were robust across separate models for children with 2+ or 3+ administrations and a full hierarchical model.The hierarchical model nested children and words within datasets, with dataset-specific mean ξ and κ values.
  • Robustness Across Model and Data Settings: The pooled model partially shrank dataset-level spread estimates toward typical cross-dataset values for datasets with few children.Datasets with many children remained essentially unconstrained in their estimated spread.
  • Sensitivity to Data-Quality Exclusions: Across four data-quality filter settings, κ estimates showed no major changes in the English and Norwegian datasets affected by the filter.The settings disabled the filter, used loose and reported thresholds, or applied a tight threshold.
  • Linear vs. Logarithmic Age: The linear model was dispreferred across all five datasets relative to the logarithmic-age M3 model.Table 7 reports the LOO ELPD advantage of log-age M3 over its linear-age counterpart, where positive values favor log-age.

2PL Model Comparison

Refitting M3 with a per-word discrimination parameter produced a better-fitting 2PL model than the Rasch 1PL across every language. Despite improved fit, substantive model results were unchanged, and low difficulty–discrimination correlations provided little support for inflated 1PL acceleration estimates.

  • 2PL Model Comparison: The comparison refit M3 by adding a per-word discrimination parameter to the Rasch model, which otherwise assigns equal discrimination to all words.
  • 2PL Model Comparison: Difficulty–discrimination correlations were small to modest, offering little support for the hypothesis that harder words substantially inflate 1PL estimates of κ.
  • 2PL Model Comparison: The 2PL fit better than the 1PL for every language, with substantial LOO ELPD differences, while other model results remained unchanged.The refit added a per-word discrimination parameter to M3. Estimates of κ increased somewhat, but κ values are not directly comparable across item models because abilities lack a common standard scale.
Loading 2608.17120v1…