Source-linked AI summary
Testing the Predictions of Surprisal Theory in 11 Languages
Ethan Gotlieb Wilcox, Tiago Pimentel, Clara Meister, Ryan Cotterell, Roger P. Levy
TL;DR
Prior evidence for surprisal theory has concentrated on English reading, leaving its crosslinguistic generality unresolved. This paper analyzes eye-tracking data from eleven languages using language-model estimates of surprisal and contextual entropy to test three theoretical predictions. All three predictions are borne out across the languages studied, although the sample remains concentrated in high-resource, industrialized settings.
Problem
Previous studies predominantly examine English reading, and no systematic crosslinguistic analysis had tested how broadly surprisal theory generalizes.
Method
Using controlled eye-tracking materials in eleven languages across five language families, the study estimates surprisal and contextual entropy from multilingual and monolingual language models and tests three predictions with regression models.
Results
All three predictions are supported in every language tested: surprisal predicts reading times, contextual entropy is predictive, and the surprisal–reading-time relationship is linear.
Takeaways & Limitations
The findings provide a robust crosslinguistic link between information-theoretic quantities and incremental language processing.
Takeaways & Limitations
The study overrepresents Indo-European languages and high-resource languages from industrialized societies, limiting coverage of the world’s languages.
Abstract
from arXiv · showhide
A fundamental result in psycholinguistics is that less predictable words take a longer time to process. One theoretical explanation for this finding is Surprisal Theory (Hale, 2001; Levy, 2008), which quantifies a word's predictability as its surprisal, i.e. its negative log-probability given a context. While evidence supporting the predictions of Surprisal Theory have been replicated widely, most have focused on a very narrow slice of data: native English speakers reading English texts. Indeed, no comprehensive multilingual analysis exists. We address this gap in the current literature by investigating the relationship between surprisal and reading times in eleven different languages, distributed across five language families. Deriving estimates from language models trained on monolingual and multilingual corpora, we test three predictions associated with surprisal theory: (i) whether surprisal is predictive of reading times; (ii) whether expected surprisal, i.e. contextual entropy, is predictive of reading times; (iii) and whether the linking function between surprisal and reading times is linear. We find that all three predictions are borne out crosslinguistically. By focusing on a more diverse set of languages, we argue that these results offer the most robust link to-date between information theory and incremental language processing across languages.
1 Introduction
The paper tests three predictions of surprisal theory across eleven languages from five language families, addressing the field’s predominantly English-focused evidence base. Using multilingual and monolingual language models with eye-tracking data, it finds support for surprisal, contextual entropy, and a linear surprisal–reading-time relationship.
- Surprisal quantifies a word’s predictability as its negative log-probability given preceding context and correlates with reading-time-based processing effort.
- Most prior studies examine English reading, leaving the crosslinguistic generality of surprisal theory systematically untested.
- The study tests whether surprisal and contextual entropy predict reading times and whether their linking function is linear.
- The analysis uses controlled eye-tracking materials in eleven languages across five language families, with multilingual and monolingual autoregressive language models estimating the predictors.The monolingual models include large- and small-corpus conditions, with the small corpora held to approximately 30 million words across languages.
- Regression models add surprisal or contextual entropy to baseline predictors and assess predictive power through held-out log-likelihood.
- All three predictions are supported crosslinguistically: surprisal improves prediction in every language, contextual entropy improves it in most, and linear models perform as well as more complex alternatives.
2 Psycholinguistic Predictive Power
The paper frames reading time as an eye-tracking measure of word-level processing difficulty and compares regression models that differ in predictors and functional form. Surprisal and contextual entropy are estimated from language models and evaluated by held-out predictive likelihood.
- Reading time is the duration readers visually attend to a word and is treated as a measure of processing difficulty.
- The analysis represents each word in context with predictor variables and fits a regression model to estimate its reading-time distribution.Reading times are treated as continuous, so the regression model is a probability density.
- Different processing theories are compared by fitting models with different predictors or architectures and evaluating held-out-data log-likelihood.Higher held-out log-likelihood indicates better predictive performance.
- The target model adds a predictor of interest to baseline predictors, while the baseline model contains only the baseline predictors.
- The study compares models using average by-word delta log-likelihood, where positive values indicate predictive power beyond the baseline.A zero delta indicates no robust relationship or no adequate approximation by the model class; negative values indicate overfitting.
- Surprisal is estimated from an autoregressive language model because the true conditional word distribution is unavailable.
- Contextual entropy is the expected surprisal of the next word given the preceding context and is likewise estimated with an autoregressive language model.
3 Experimental Setup
The experiment combines crosslinguistic eye-tracking data with monolingual and multilingual language models to estimate surprisal and contextual entropy. Regression analyses compare these predictors with reading-time baselines across languages, contexts, and model-training scales.
- 3.1 Dataset: MECO provides comparable eye-tracking data from L1 readers across thirteen languages and five language families, with Norwegian and Estonian excluded from multilingual-model analyses.The corpus contains 12 simplified Wikipedia-style articles translated to preserve content across languages.
- 3.1 Dataset: Reading-time analyses use first fixation, gaze duration, and total fixation, with gaze duration emphasized as the measure most closely associated with first-pass processing difficulty.Skipped first-pass words receive a reading time of zero and remain in the analysis.
- 3.2 Language Models: Surprisal and contextual entropy estimates come from multilingual and monolingual autoregressive language models, including monolingual models trained on all available data or approximately 30 million tokens.The monolingual models use Wiki40B data, language-specific 32k UnigramLM tokenizers, and decoder-only transformers.
- 3.2 Language Models: The study estimates surprisal in short and longer contexts to examine whether context-window size biases psycholinguistic predictions.The context-length manipulation is motivated by the possibility that short contexts shift probability mass away from low-frequency words.
- 3.2 Language Models: Model psychological plausibility is considered relative to human linguistic exposure, while the authors argue that training-data size alone does not determine human-like predictions.The compared training scales range from young-child-like exposure to multiple human lifetimes of language data.
- 3 Experimental Setup: Regression models predict word reading times from current and preceding-word predictors, use 10-fold cross-validation, and assess target-versus-baseline differences with paired permutation tests.Baseline predictors include word length and negative log frequency for the current and two preceding words.
4 Results
Across languages, surprisal improves reading-time prediction, while contextual entropy generally helps when added but hurts when substituted. Word-level coefficients show consistent effects for surprisal and length, with cross-language variation and no significant link between mGPT perplexity and predictive gains.
- 4.1 Surprisal: Surprisal significantly improves reading-time prediction across languages, with aggregate effects greater than zero in all tested cases.For gaze duration with mGPT, predictive-power gains fall between 0.012 and 0.040 across languages.
- 4.2 Contextual Entropy: Adding contextual entropy generally improves prediction, whereas replacing surprisal with entropy usually reduces predictive power.For mGPT gaze duration, adding entropy yields positive gains in 8/11 languages, while replacement yields negative gains in 7/11.
- 4.2 Contextual Entropy: Contextual entropy has a weak but consistent association with reading times across languages, consistent with readers pre-planning processing based on expected surprisal.The aggregate add model is significantly positive for all three reading-time measures.
- 4.3 Variation Across Languages: mGPT perplexity does not significantly correlate with predictive gains across languages or language families.The by-language correlation is negative but nonsignificant (ρ = −0.186, p = 0.6), and the analysis includes mGPT only.
- 4.4 Model Coefficients: Surprisal produces a consistent 2-4 ms/bit effect for the current word, smaller 0-2 ms/bit effects for the preceding word, and no obvious effect two words earlier.The current-word effect is smallest in Hebrew and larger in Dutch, Russian, Greek, and Italian.
- 4.4 Model Coefficients: Word length has positive effects for the current word and consistent negative effects for the preceding word across languages.The preceding-word pattern may reflect word skipping after a long word, which produces zero reading time in the analysis.
5 Surprisal–RT Linking Function
The analyses support a linear relationship between surprisal and reading-time slowdown. Across languages, nonlinear models offered no consistent predictive advantage over models constraining this relationship to be linear.
- Visualizing the Link with GAMs: The analysis used GAMs with matched baseline effects and 10-fold cross-validation to compare flexible and linear surprisal effects.The models included surprisal, frequency, and length predictors, with baseline specifications held comparable for this analysis.
- Visualizing the Link with GAMs: GAMs recovered an approximately linear surprisal–reading-time relationship across languages and contexts.The nonlinear fits sometimes overlapped directly with the linear control fits.
- Testing Linearity: The study compared linear and nonlinear GAMs against a shared baseline to test whether nonlinear surprisal effects improved predictive power.A nonlinear advantage would support nonlinearity; a consistently null difference would support linearity.
- Testing Linearity: Permutation tests found no significant difference between linear and nonlinear models for any tested language or model at α = 0.05.The comparison used differences in predictive power relative to the shared baseline.
- Testing Linearity: The results support the linear link hypothesis, despite prior studies reporting superlinear and sublinear alternatives.The authors note that the tested results were reported for gaze duration.
6 Discussion
The discussion reports stable surprisal effects across diverse languages and model types, while identifying limits on the study’s linguistic and theoretical generality. It also highlights weaker previous-word effects and argues that the linear relationship does not support a universal channel-capacity account.
- Limitations and Future Directions: The language sample includes multiple word orders, headedness patterns, and morphological types, but Indo-European languages remain overrepresented.Each non-Indo-European family is represented by a single language.
- Implications of Psycholinguistic Theories: For gaze duration and mGPT, surprisal Δ values ranged from 0.012 to 0.040 across languages.Across languages and models, surprisal showed relatively little variance in predictive power.
- Implications of Psycholinguistic Theories: The surprisal effect was 2–4 ms/bit in every language, encompassing the earlier English estimate of 3.75 ms/bit.The earlier estimate used n-gram-derived surprisal, which is generally higher than estimates from large neural language models.
- Implications of Psycholinguistic Theories: The authors argue that the linear surprisal–reading-time relationship does not support a universal channel-capacity account requiring a superlinear link.They also report no consistent superlinear pattern for larger models or longer contextual windows.
- Implications of Psycholinguistic Theories: Previous-word surprisal effects were weaker in this study, ranging from 0–2 ms/bit.The authors relate this pattern to strongly incremental processing measures and prior eye-tracking findings.
- Limitations and Future Directions: The study covers high-resource written languages and participants from industrialized societies, limiting its coverage of global linguistic variation.The method also requires a large written corpus, which may be unavailable for many lower-resource languages.
7 Conclusion
The paper evaluates three surprisal-theory predictions across eleven languages and five language families. All three predictions were supported in every language tested, yielding a robust crosslinguistic link between information-theoretic quantities and incremental processing.
- Conclusion: The study tested surprisal, contextual entropy, and linear-link hypotheses using controlled eye-tracking materials in eleven languages across five language families.The three hypotheses concern prediction of reading times, prediction by contextual entropy, and linearity of the surprisal–reading-time relationship.
- Conclusion: All three predictions were borne out in every language tested, showing exceptionally strong crosslinguistic stability.The conclusion characterizes these findings as the most robust link between information-theoretic quantities and incremental processing.
A Regression Modeling Details
The paper uses mixed-effects regressions to test surprisal and contextual-entropy effects on reading time, with language-specific and aggregate models. It also compares nonlinear and linear GAMs for the surprisal–reading-time relationship while controlling for lexical and preceding-word variables.
- Regression variables distinguish the current word from the previous two words through prev and prev2 prefixes.
- The individual-language surprisal model predicts reading time from current, previous, and two-previous-word surprisal alongside frequency and length controls.
- Aggregate surprisal models add language-varying effects for surprisal, frequency, and length predictors.
- Contextual-entropy tests compare models that include surprisal, length, and frequency baselines with models adding or replacing predictors with entropy terms.
- Nonlinear GAMs use spline smooths for surprisal and previous surprisal, whereas linear GAMs model those effects directly as linear terms.The nonlinear smooth uses up to six basis functions, allowing relatively simple non-linear curves.
B Surprisal versus RT for wt−1
The relationship between reading time and the surprisal of the previous word varies across languages. Several languages show roughly linear increases, while others show nonlinear, negative, or difficult-to-interpret relationships.
- Earlier English work found a previous-word surprisal effect about as strong as the current-word surprisal effect.
- The analysis uses gaze duration and compares nonlinear GAMs with linear control GAMs, including bootstrapped 95% confidence intervals.
- Previous-word surprisal effects are roughly linear and increasing in English, Italian, Korean, Russian, Greek, and Spanish.
- Dutch and Turkish show roughly increasing but visually nonlinear relationships between reading time and previous-word surprisal.
- Finnish, German, and Hebrew show negative or difficult-to-interpret relationships, aligning with weak or sometimes negative coefficients.