Source-linked AI summary
Do Language Models Know What Not to Say? Causal Evidence for Statistical Preemption in LLMs
Dongxin Guo, Jikun Wu, Siu Ming Yiu
TL;DR
The paper investigates how learners acquire knowledge of unacceptable forms without explicit negative evidence and whether LLMs distinguish statistical preemption from entrenchment. Across four experiments and three English construction types, it combines corpus-based surprisal, human validation, scaling analysis, and controlled fine-tuning. LLMs reproduce human-like preemption sensitivity, including causal shifts from competing-form frequency, while the evidence remains limited to English and does not fully resolve alternative interpretations.
Problem
The paper addresses how learners reject structurally possible forms without explicit negative evidence and whether distributional competition, rather than overall verb frequency, explains this knowledge.
Method
Across four experiments, the study analyzes 120 English verb–construction pairings using 14 LLMs, corpus-derived preemption and entrenchment measures, human datasets, scaling analysis, and controlled fine-tuning interventions.
Results
LLM surprisal correlates with human acceptability at r = 0.79, and manipulating conventional-form frequency shifts preemption behavior with an Amplified effect of +0.73 versus −0.29 for Reverse.
Takeaways & Limitations
The findings support distributional competition as a mechanism through which neural language models acquire negative linguistic knowledge resembling statistical preemption.
Takeaways & Limitations
Claims are restricted to English and three construction types, and the Amplified–Reverse asymmetry remains compatible with pre-existing corpus imbalance and embedding-neighborhood effects.
Abstract
from arXiv · showhide
How do learners acquire knowledge of what is unacceptable without negative evidence? Construction Grammar proposes statistical preemption: exposure to a conventional form (e.g., "donated the books to the library") preempts structurally possible but unattested alternatives ("*donated the library the books"). We present a computational study that, for the first time, directly dissociates statistical preemption from the competing entrenchment hypothesis in large language models within a single converging design. Across four experiments spanning 120 English verb-construction pairings (dative, causative, locative), we show that (1) LLM surprisal patterns correlate strongly with human acceptability judgments ($r = 0.79$), validated against three independent behavioral datasets; (2) these patterns are driven by competing-form frequency rather than overall verb frequency, confirmed by non-circular partial correlations; (3) preemption sensitivity scales as a power law with model size; and (4) a controlled fine-tuning intervention causally demonstrates that manipulating competing-form frequencies shifts preemption behavior in the predicted direction, with reverse-direction controls ruling out frequency-sensitivity confounds. These results provide converging evidence that neural language models acquire negative linguistic knowledge through distributional competition, the core mechanism posited by Construction Grammar.
1 Introduction
The paper asks how learners acquire knowledge of unacceptable forms without explicit negative evidence, focusing on statistical preemption versus entrenchment. It tests whether LLM distributional learning reproduces human-like preemption across multiple experiments and construction types.
- Motivation: Baker’s Paradox asks how learners reject structurally possible forms such as “*She donated the library the books” despite lacking explicit negative evidence.The contrast arises because semantically similar verbs like give allow the double-object construction.
- Theoretical framework: Statistical preemption proposes that frequent conventional alternatives inhibit functionally equivalent unattested forms, unlike entrenchment based on overall verb frequency.The competing accounts make different predictions about how constructional distributions affect acceptability.
- Research question: The study tests whether LLM surprisal captures preemption’s distributional signature and aligns item-by-item with human acceptability judgments.The motivation is to assess whether distributional learning alone can yield negative linguistic knowledge without explicit grammatical instruction.
- Experimental program: Four experiments examine preemption effects, disentangle competing-form frequency from entrenchment, model scaling, and manipulate training frequencies causally.The design spans dative, causative, and locative alternations and includes non-circular human validation and a reverse-direction control.
2 Background
Construction Grammar treats preemption as competition from a conventional alternative that constrains overgeneralization gradually. Prior behavioral and computational work motivates a design that directly separates competing-form effects from entrenchment using known training distributions and human comparisons.
- Construction Grammar: Statistical preemption predicts that stronger conventional competition produces greater gradient unacceptability for an alternative construction.The mechanism balances productive coverage against the inhibitory force of established alternatives.
- Behavioral evidence: Behavioral studies support competing-form frequency effects, while entrenchment and preemption have often been difficult to separate because their predictors are correlated.Winkler et al. provided a +Competing/−Competing dissociation, whereas other work reported mixed or jointly reliable effects.
- Empirical domain: The English dative alternation provides a broad test bed because verbs range from freely alternating to strongly constructionally restricted.Restricted examples include donate, explain, and whisper, whereas cost and fine resist the prepositional alternative.
- LLMs and Construction Grammar: LLMs offer a complementary test because their known training distributions permit independent computation of preemption and entrenchment scores and comparison with human judgments.Prior work has used surprisal and controlled-input training to study constructional learning and human alignment.
- Contribution: The study’s methodological contribution combines multi-construction scope, non-circular human validation, scaling analysis, and a reverse-direction causal control.This combination is positioned against closely related work summarized in the feature comparison.
3 Experimental Design
The experiments cover 120 English verb–construction items across dative, causative, and locative alternations, using controlled sentence frames, multiple model families, corpus-derived measures, and complementary statistical tests. Human comparisons cover dative and causative data, while locative results use corpus-based predictions because no comparably large item-level dataset exists.
- Stimuli: The stimulus set contains 120 items organized by preemption strength across dative, causative, and locative constructions.Datives include 80 verbs divided into strong, weak, and no-preemption categories; causative and locative sets contain 20 verbs each.
- Stimuli: Each verb appears in five matched sentence frames controlling length, subject animacy, object definiteness, and tense.The design aims to align model surprisal differences with stable verb-level construction preferences.
- Language models: The study evaluates 14 base models from four families, with Pythia supplying controlled scaling and OLMo enabling direct corpus verification.Instruction-tuned models are excluded because tuning can shift surprisal patterns in ways that complicate psycholinguistic interpretation.
- Human benchmarks: Human comparisons use DAIS, Robenalt and Goldberg, and Tachihara and Goldberg for dative and causative alternations, but locatives lack a comparably large item-level behavioral dataset.Locative results are therefore compared with corpus-based classifications and effect-size predictions rather than direct item-level judgments.
- Analysis: The analysis combines effect-size tests, Pearson correlations with bootstrap intervals, mixed-effects regressions, and partial correlations with FDR correction across 94 tests.Surprisal is averaged across five frames for each verb, and corpus parsing uses construction-specific dependency templates.
- Measures: Preemption is operationalized from conventional-versus-unconventional construction frequencies, while entrenchment is total log verb frequency.The preemption score is a distributional proxy for functional competition, and the paper explicitly notes this theoretical gap.
4 Experiment 1: Preemption Effects
LLMs show the predicted graded preemption pattern, with larger surprisal differences for strongly preempted verbs, and their item-level predictions correlate strongly with human acceptability. The ordering of effects across dative, causative, and locative alternations also matches human evidence.
- 4.2 Results: Group Differences: 2.41 bits/word for strongly preempted dative verbs, versus 1.12 for weakly preempted and 0.33 for non-preempted verbs in LLaMA-2 7B.The strong-versus-no-preemption contrast is highly significant, t(52) = 9.87, p < .001, d = 2.87.
- 4.2 Results: Group Differences: The predicted graded pattern appears across all 14 models, not only the representative LLaMA-2 7B results.Complete results cover all three constructions in the appendix.
- 4.3 Results: Correlation with Human Judgments: r = 0.79 [0.69, 0.86] between LLaMA-2 7B surprisal differentials and DAIS human acceptability across 80 verbs.All models above 1B parameters exceed r = 0.70, with replication at r = 0.74 and r = 0.76 in two additional behavioral datasets.
- Cross-construction results: The preemption gradient extends beyond datives: causatives show ∆S = 2.17 versus 0.32 for non-preempted items, while locatives show a smaller d = 1.42.The effect-size ordering is dative d = 2.87 > causative d = 2.34 > locative d = 1.42, matching human ordering.
5 Experiment 2: Preemption vs. Entrenchment
Experiment 2 separates statistical preemption from entrenchment by comparing frequency-matched verbs with and without a dominant competing construction. Across LLMs and human-linked analyses, competing-form frequency explains substantially more variance than overall verb frequency.
- Design: The design classifies +Competing verbs by a dominant functionally equivalent alternative and matches them to −Competing verbs on overall frequency and four other potential confounds.The groups were matched on log frequency, verb-class entropy, morphological complexity, register distribution, and concreteness; no matched variable differed significantly.
- Results: For LLaMA-2 7B, frequency-matched +Competing verbs produced larger surprisal differences than –Competing verbs (∆S = 2.36 vs. 0.91; d = 1.91).The effect was robust across models, with d ranging from 1.43 to 2.18.
- Regression Analysis: Preemption was the dominant predictor of ∆S (β = 3.41), whereas entrenchment contributed modestly (β = 0.19).Partial correlations were rpartial = 0.72 for preemption controlling entrenchment and rpartial = 0.24 for the reverse.
- Non-Circular Test: Corpus-derived preemption predicted human acceptability beyond entrenchment (rpartial = 0.58 for DAIS), while the reverse association was weak and nonsignificant (rpartial = 0.12).The same asymmetry appeared in R&G ratings, with rpartial = 0.52 for preemption and 0.08 for entrenchment.
- Non-Circular Test: The triangulated analyses decompose the overall LLM–human correspondence (r = 0.79) into asymmetric links favoring preemption rather than entrenchment.Additional raw-frequency, n-gram, and primacy-of-human-data controls converged on the same conclusion.
6 Experiment 3: Scaling Behavior
Experiment 3 tests whether preemption sensitivity changes systematically with model size. Across the Pythia suite, sensitivity increased monotonically and followed a well-fitting sublinear power law, with similar behavior across architectures.
- Scaling Behavior: Preemption sensitivity increased monotonically across Pythia models from 160M to 12B parameters and followed a power law with b = 0.092.The fit had adjusted R2 = 0.993, indicating diminishing returns rather than a sudden phase transition.
- Scaling Behavior: Cross-architecture models clustered near the Pythia trend, supporting generalizability beyond one model family.Alternative functional forms fit worse, and the qualitative scaling result survived leave-one-out checks across the six Pythia models.
7 Experiment 4: Causal Intervention
Experiment 4 uses controlled fine-tuning to manipulate competing-form frequencies and test whether preemption behavior changes causally. Increasing the conventional competing form produced a larger effect than reversing the manipulation, while non-target verbs remained unchanged.
- Design: The intervention compares amplified conventional-form frequency, balanced dative frequencies, and amplified unconventional-form frequency for the same target verbs.The Amplified and Reverse conditions each multiply one form’s frequency by three, while Attenuated equalizes both forms.
- Results: The Amplified condition increased ∆S by 0.73 ± 0.07 bits, significantly exceeding the Reverse effect of −0.29 (p < .001).The asymmetry matches preemption theory’s prediction and is not explained by simple frequency sensitivity alone.
- Results: The Attenuated condition decreased ∆S by 0.43 ± 0.05 bits, whereas control verbs showed no change.The intervention was replicated across five random seeds.
- Results: Non-target verbs showed no systematic change (∆∆S = +0.02 ± 0.05, p = .71), supporting verb-specificity.The design also reports that the Amplified–Reverse asymmetry correlates with change in preemption ratio rather than raw frequency change.
- Limitations: The authors acknowledge that pre-existing corpus imbalance and embedding-space neighborhood effects could partly contribute to the Amplified–Reverse asymmetry.Verb-specificity addresses but does not fully exclude subtler representational effects.
8 Implications for Linguistic Theory
The study supports statistical preemption as learnable from positive distributional evidence, while leaving open whether the effect reflects usage-based competition or structured regularities. Its claims remain bounded by English-only testing and a distributional proxy for functional equivalence.
- 8 Implications for Linguistic Theory: LLMs’ competing-form sensitivity supports Goldberg’s claim that preemption is learnable from positive evidence alone, without innate semantic verb-class constraints.The non-circular partial correlations show that the same distributional variable predicts model behavior and human judgments.
- 8.2 The Formal–Functional Divide: The distributional proxy does not directly capture functional equivalence, so pragmatic constraints, register effects, or structured regularities could produce similar preemption scores.Excluding eight plausibly register-driven verbs strengthened the effect, but the authors do not treat this as adjudicating usage-based versus formal accounts.
- 8.3 Cross-Linguistic Predictions: The study tests only English, although preemption may operate over morphological alternations in agglutinative languages and different constructions in isolating languages.Cross-linguistic testing is identified as the most critical next step.
- 8.4 Relationship to Semantic Verb-Class Accounts: Frequency-matched verbs from similar semantic classes show different ∆S depending on whether a competing form exists, constraining purely semantic-class accounts.The study cannot fully rule out implicit semantic learning, which would require testing narrower verb classes with differing preemption strength.
9 Discussion
Across four experiments, LLMs reproduce the distributional signature of statistical preemption, with behavior causally shifted by conventional-competitor frequency. The study frames this as evidence that models can help explain why forms are unacceptable, not merely detect that they are.
- 9 Discussion: LLMs reproduce statistical preemption and are causally modulated by the frequency of conventional competitors across the study’s four experiments.Non-circular corpus-to-human partial correlations indicate that conventional alternatives, rather than exposure alone, predict verb restrictions.
- 9 Discussion: The study shifts the question from whether language models register unacceptability to why, locating the answer in competition rather than exposure alone.This positions preemption sensitivity within broader work treating LLMs as scientific instruments.
- 9 Discussion: The paper concludes that neural language models develop the preemption sensitivity shaping human acceptability in English, while typological generalization remains open.The conclusion identifies cross-linguistic extension as the central unresolved question.
Limitations
The study’s interpretation is limited by its English-only scope, incomplete human validation for locatives, proxy measures of functional competition, and a fine-tuning intervention that does not recreate development. Additional concerns include unresolved mechanisms, alternative interpretations of the reverse asymmetry, limited causal scale, and reliance on existing datasets and base models.
- English-only scope: All claims are restricted to English and three construction types, so cross-linguistic generalization remains untested.The paper does not test its appendix predictions for typologically diverse languages.
- Locative human-data asymmetry: Locative results lack item-level human validation because large matched behavioral datasets exist only for the dative and causative alternations.They are evaluated against corpus classifications and effect-size predictions instead.
- Distributional proxy for functional competition: Corpus-based preemption scores are distributional proxies for functional competition, and register exclusion plus group dissociation do not eliminate that theoretical gap.Pragmatic constraints, register effects, or structured regularities could also generate high scores.
- Fine-tuning does not reconstruct developmental learning: Fine-tuning demonstrates sensitivity to competing-form frequency but does not recreate the developmental trajectory through which preemption preferences are originally acquired.Controlled-rearing designs from initialization are described as stronger developmental tests.
- Scale and alternative interpretations: The causal intervention uses GPT-2 124M with 20 verbs, and replication with larger models and broader samples would strengthen the claims.The reverse asymmetry is also compatible with pre-existing corpus imbalance and cannot be fully interpreted as uniquely preemptive.
- Mechanism and data limitations: The study does not probe the internal mechanism implementing preemption and relies on existing human datasets and base models rather than matched judgments or instruction-tuned models.These choices leave mechanistic and model-family generalization questions unresolved.
D.1 Robustness: Low-Collinearity Subset
Robustness analyses show that the preemption effect survives low-collinearity filtering, alternative surprisal measures, corpus-model controls, classification validation, and perturbations to parsing thresholds. The construction-specific pipeline uses dependency patterns to identify the relevant alternations.
- D.1 Robustness: Low-Collinearity Subset: β = 3.18 for preemption remains significant in the low-collinearity subset, while entrenchment becomes marginal.This stability supports a distinct preemption contribution when PREEMPT and ENTRENCH correlations are below 0.3.
- Robustness controls: r = 0.77 with DAIS under SLOR normalization and r = 0.75 with critical-region surprisal, closely matching the main human-correlation result.The reported correspondence is therefore not tied to one surprisal formulation.
- Corpus-model circularity controls: R2 = 0.68 for PREEMPT versus R2 = 0.41 for raw co-occurrence, while the 5-gram baseline reaches r = 0.41 versus r = 0.79 for transformer LLMs.These controls distinguish the preemption measure from raw frequency and surface co-occurrence alone.
- E A Priori Classification Validation: 116/120 verbs receive matching BNC- and Dolma-based classifications, with Cohen’s κ = 0.94, after classifications were finalized before ∆S computation.This supports the independence and reliability of the a priori classification procedure.
- G.2 Construction-Specific Templates: The pipeline identifies dative, causative, and locative frames using construction-specific dependency patterns and excludes third-frame locative instances from PREEMPT counts.Causative periphrastic frames are separately detected and excluded from unconventional-frame counts.
G.4 Validation Method and Precision
The validation pipeline combined manual annotation, parser perturbation checks, and behavioral comparisons to assess the reliability and scope of preemption scores. Results were broadly stable, with strongest agreement for datives and weaker correspondence for causatives and locatives.
- Validation: 96% dative, 93% causative, and 92% locative pipeline precision were obtained against adjudicated annotations.Inter-annotator agreement was κ = 0.94, 0.91, and 0.89, respectively.
- Validation: 87% of rejected sentences were genuine non-matches, while the remaining 13% were predominantly parser errors on long or coordinated sentences.
- Robustness: Parser-threshold and matcher perturbations preserved per-verb preemption scores, with correlations of r ≥0.93 for datives, r ≥0.89 for causatives, and r ≥0.85 for locatives.The analyses doubled or halved the confidence threshold and replaced strict dependency matching with a permissive POS-pattern matcher.
- Discrepancies: Excluding four verbs with register restrictions or polysemy increased the LLM–human correlation from r = 0.79 to r = 0.83 and strengthened the preemption–entrenchment dissociation.Several low-frequency verbs showed ∆S > 1.5 despite near-chance human DAIS bias scores.
- Behavioral validation: LLM–human agreement was strongest for datives (RMSE = 0.31), intermediate for causatives (RMSE = 0.38), and weakest for locatives (RMSE = 0.47).This ordering mirrors the strength of preemption effects in human data and may reflect differences in corpus frequency and alternation regularity.