Source-linked AI summary
From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish
Francisco Portillo López
TL;DR
The paper asks whether contextual surprisal contributes to explaining children’s word acquisition as much as it predicts adult reading times. Across two Spanish studies, it compares surprisal with frequency in AoA and eye-tracking analyses, finding a stronger surprisal association in adult reading than in acquisition.
Problem
It is unclear whether frequency and contextual surprisal reflect a common mechanism or become behaviorally relevant at different stages of language development.
Method
Two corpus-based Spanish studies model AoA for 225 nouns and adult fixation durations using frequency measures and LLM-derived surprisal.
Results
z = 3.63, p < .001: the surprisal–behavior association was stronger in adult reading than in acquisition.
Takeaways & Limitations
Cumulative lexical exposure is particularly informative about early acquisition timing, whereas surprisal captures processing difficulty in an established linguistic system.
Takeaways & Limitations
The amount of context available for surprisal estimation differs across studies, so the cross-domain contrast could reflect richer context in Study 2.
Abstract
from arXiv · showhide
Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it contribute as much to explaining when children acquire individual words? Frequency reflects a learner's cumulative exposure to a word, whereas surprisal reflects how predictable a single occurrence is given its context. We investigate this question across two corpus-based studies of Spanish. In Study 1, we modeled age of acquisition (AoA) for 225 Spanish nouns using lexical frequency and contextual diversity from child-directed speech, plus surprisal from three language models differing in architecture and training language (BETO, BERTIN, mGPT). Frequency strongly predicted AoA (r=-.597, p<.001); surprisal added little beyond frequency and word length, including in a naturalistic-context analysis. In Study 2, we modeled adult fixation durations in the Chilean Spanish subsample of the Multilingual Eye-movement Corpus (MECO Wave 2), using mGPT surprisal alongside two independent frequency measures. Surprisal robustly predicted longer fixation durations after controlling for frequency and word length, consistent across both frequency sources. A matched word-type-level comparison showed the surprisal-behavior association was stronger in reading than in acquisition (z=3.63, p<.001). The findings suggest cumulative lexical exposure and contextual predictability play different roles across the language trajectory: frequency is particularly informative about when early lexical representations are acquired, whereas surprisal captures moment-to-moment processing difficulty in an already-established linguistic system. We discuss this pattern in relation to usage-based and entrenchment-based accounts of lexical development and to the evaluation of language models as models of human language behavior.
1 Introduction
The introduction distinguishes cumulative lexical exposure from contextual predictability and asks whether surprisal contributes similarly to lexical acquisition and adult reading. The paper addresses this question in Spanish across developmental and online-processing measures.
- Research question: Age of acquisition concerns when a lexical item becomes established in a child’s productive vocabulary, whereas surprisal concerns processing difficulty for a particular occurrence.These measures operate over different timescales: acquisition accumulates across encounters, while surprisal reflects predictability in context.
- Research question: Frequency measures cumulative experience with a word, while surprisal measures the predictability of an individual occurrence given its context.The two measures can therefore make different predictions about developmental learning and online processing.
- Theoretical perspectives: Usage-based and entrenchment-based accounts predict that repeated encounters strengthen lexical representations and make cumulative frequency relevant to acquisition timing.Error-driven accounts instead allow poorly predicted contexts to produce comparatively more learning, so a null surprisal–AoA relation is not theoretically predetermined.
- Background: Prior work links frequency and related distributional measures to early word acquisition and LLM-derived surprisal to adult reading behavior.The introduction frames the unresolved issue as whether these findings reflect one mechanism or stage-specific behavioral relevance.
- Present study: The study examines Spanish across two behavioral domains: developmental establishment of lexical knowledge and online processing in an already-established linguistic system.Study 1 tests surprisal for AoA beyond frequency and contextual diversity using BETO, BERTIN, and mGPT; Study 2 tests mGPT surprisal in Spanish adult reading beyond frequency and word length.
2.1 Method
Study 1 combines Spanish child vocabulary norms, child-directed speech measures, and surprisal from three language models to model AoA for Spanish nouns. It also tests standardized and naturalistic contexts.
- Data: The study models age of acquisition for 225 Spanish nouns using Spanish CDI norms from Wordbank as the dependent measure.The merged corpus-derived analyses include 224 nouns because one target lacked the required merged measures.
- Data: Child-directed speech from CHILDES supplies lexical frequency and contextual-diversity measures for the target nouns.Frequency counts occurrences, while contextual diversity counts distinct utterances or contexts containing a noun.
- Predictors: BETO, BERTIN, and mGPT provide surprisal estimates while differing in architecture and training data.Surprisal is calculated as the negative log probability assigned to the target word given preceding context.
- Predictors: The primary analysis embeds nouns in standardized carrier frames to control contextual predictability across items.An additional mGPT analysis averages surprisal across up to ten valid naturalistic child-directed-speech contexts per noun.
- Analysis: Regression models first predict AoA from lexical frequency and word length, then test the incremental contribution of mGPT surprisal.Natural-context analyses include 224 nouns, with a robustness analysis restricted to 154 nouns having at least five valid contexts.
2.2 Results
Frequency was strongly associated with earlier AoA, whereas surprisal added little independent explanatory value after frequency and word length were controlled. Naturalistic contexts produced a small raw association but no reliable incremental contribution.
- Frequency and contextual diversity: ρ = −.597, p < .001: child-directed-speech frequency was associated with earlier acquisition of Spanish nouns.Words occurring more often in the input tended to be acquired earlier.
- Frequency and contextual diversity: Contextual diversity showed a similar AoA relationship, but its independent contribution was substantially reduced after controlling for frequency.The pattern identifies cumulative exposure as the primary distributional correlate of acquisition timing in these data.
- Robustness analysis: 154 nouns with at least five valid contexts: the natural-context association was no longer significant (ρ = .132, p = .104).The adjusted model produced ∆R2 = .0076, and the mGPT coefficient remained non-significant (β = −.0737, p = .197).
2.3 Interim discussion
Study 1 provides strong evidence that cumulative lexical exposure relates to acquisition timing, while contextual surprisal contributes little beyond frequency and word length in this dataset. This motivates testing whether surprisal predicts adult online reading difficulty.
- Study 1 interpretation: Words occurring more frequently in child-directed speech tended to be acquired earlier, replicating the established frequency–AoA relationship.Contextual diversity was also associated with AoA, but its contribution was substantially reduced when frequency was controlled.
- Study 1 interpretation: Standardized-context mGPT surprisal was unrelated to AoA and added virtually no explanatory value beyond frequency and word length.Naturalistic contexts produced a small raw association, but it disappeared after established distributional predictors were controlled.
- Interpretation: The findings suggest that moment-to-moment predictability provides limited information about acquisition timing once cumulative exposure is taken into account.This does not imply that children never use contextual expectations during learning.
- Theoretical implications: The pattern is compatible with usage-based and entrenchment-based accounts in which repeated encounters strengthen lexical representations in memory.The study does not adjudicate between entrenchment-based and error-driven accounts of acquisition.
- Motivation for Study 2: Study 2 tests whether mGPT surprisal predicts moment-to-moment reading difficulty in adult Spanish despite its limited acquisition contribution.This comparison evaluates whether the model captures psychologically relevant structure for adult processing rather than early lexical acquisition.
3 Study 2: Surprisal and Adult Reading Times in Spanish
Study 2 tested whether contextual surprisal explains variation in adult Spanish reading times beyond lexical frequency and word length. Using naturally occurring discourse contexts and two independent frequency norms, surprisal robustly predicted longer fixation durations.
- Data and measures: mGPT estimated each target word’s surprisal from the full preceding passage, capturing contextual predictability during continuous reading.Word-level surprisal was computed in bits, with constituent sub-word surprisals summed for multi-token words.
- Modeling approach: The mixed-effects models predicted log total fixation duration from surprisal, frequency, and word length, with crossed random intercepts for participant and word type.Parallel models used EsPal and SUBTLEX-ESP frequency norms because they were highly correlated (r = .95).
- Results: β = 0.0114 (SE = .00067, t = 17.19, p < .001) for surprisal with EsPal frequency, and β = 0.0114 (SE = .00066, t = 17.33, p < .001) with SUBTLEX-ESP.Less predictable words received longer total fixation durations after controlling for lexical frequency, word length, and crossed random effects.
- Results: β = 0.112 (95% CI [0.099, 0.125]) made surprisal the second-largest standardized lexical predictor, after word length and ahead of frequency.Word length had β = 0.292, while lexical frequency had β = −0.051; all three effects were statistically significant.
- Interim discussion: Unlike in Study 1’s acquisition analysis, contextual surprisal explained adult moment-to-moment processing variation beyond frequency and word length.The contrast was observed when surprisal was estimated from the preceding linguistic context of naturally occurring text.
4 General Discussion
Across development, frequency and surprisal show different predictive roles: frequency is informative about age of acquisition, whereas surprisal is more informative about adult reading behavior. The cross-study difference supports evaluating language models across populations and behavioral levels.
- The studies asked whether LLM-derived surprisal plays the same explanatory role in lexical acquisition and skilled adult reading.
- Frequency strongly predicted AoA, while surprisal added little independent information after frequency and word length were controlled.
- Naturalistic child-directed contexts produced a small surprisal–AoA association, but it disappeared after controlling for frequency and word length.
- In adult reading, higher surprisal predicted longer fixation durations across independent frequency measures, with surprisal ranking second among the three lexical predictors after word length.
- The authors distinguish cumulative frequency from contextual predictability, linking frequency more closely to lexical establishment and surprisal to processing established lexical knowledge.
- The interpretation is theoretical rather than a demonstrated developmental mechanism because neither study experimentally manipulated frequency.
- The matched comparison found that surprisal’s association with behavior was stronger in reading than acquisition: z = 3.63, p < .001.
- Differences in context length, Spanish variety and register, lexical distributions, and item composition constrain the cross-study contrast.
Ethics Statement
The research used publicly available, de-identified secondary data and did not involve new participant data collection. The original data collections were conducted and ethically approved by the respective corpus teams.
- The analyses used publicly available, de-identified secondary data from Wordbank CDI norms, CHILDES transcripts, and the MECO Wave 2 corpus.
- The author conducted no new data collection from human participants.
- No institutional review was required for the present analyses because the work analyzed secondary data.
- Original data collection was conducted and ethically approved by the respective corpus teams.
Funding
The research received no specific grant from any funding agency.
- The research received no specific grant from any public, commercial, or other funding agency.
Conflict of Interest Statement
The author declares no conflict of interest.
- The author declares no conflict of interest.
Open Science Statement
Both studies were exploratory, corpus-based analyses of secondary data and were not preregistered. Their statistical comparisons are therefore hypothesis-generating as well as hypothesis-testing where relevant, and the results remain provisional pending independent replication.
- Both studies were exploratory, corpus-based analyses of secondary data and were not preregistered.
- Reported statistical comparisons should be interpreted as hypothesis-generating in addition to hypothesis-testing where relevant.
- The results should be considered provisional pending independent replication.
Data Availability
The study materials and analysis resources are publicly available through established repositories, including Wordbank, TalkBank, and the Open Science Framework.
- CDI norms are publicly available through Wordbank.
- CHILDES corpora are publicly available through TalkBank.
- MECO Wave 2 data, release 2.0, are publicly available through the Open Science Framework.
- Analysis code for both studies is available through the cited Open Science Framework resource.