Source-linked AI summary
Which Forms of Caregiver Feedback Support Grammar Learning? A Reinforcement-Learning Study of Child-Like Language Models
Jing Liu, Marianne Schweitzer, Abdellah Fourtassi
TL;DR
The study asks how different caregiver-feedback forms contribute to children’s grammatical learning, a question difficult to isolate in naturalistic interaction. It uses child-like language models with pretraining and reward-based fine-tuning to compare feedback types. Structural alignment most consistently improves grammaticality in generation, while other feedback forms show weaker or negative grammar effects.
Problem
The causal effects of different caregiver-feedback types on children’s grammatical learning are difficult to isolate because these signals co-occur in naturalistic interaction.
Method
Small child-like language models are pretrained on child-directed language and then reinforcement-fine-tuned with reward models representing communicative feedback, structural alignment, semantic contingency, and affective feedback.
Results
Structural alignment systematically improves grammaticality in generation, communicative feedback yields smaller gains, and semantic contingency and affective feedback do not improve grammaticality.
Takeaways & Limitations
Different caregiver-feedback types push language generation in different directions and may support complementary aspects of language learning beyond grammar.
Takeaways & Limitations
The study captures only verbal feedback and uses a single GPT-2-style architecture, limiting coverage of multimodal learning and architectural generality.
Abstract
from arXiv · showhide
Social interaction is central to children's language learning, but the effects of different forms of caregiver feedback are difficult to isolate in naturalistic data. We use child-like language models as controlled learners to test which forms of feedback support grammatical development. Small GPT-2-style models are pretrained on child-directed language from CHILDES, then fine-tuned with reinforcement learning using reward models trained to capture four feedback types: communicative feedback, structural alignment, semantic contingency, and affective feedback. Reward fine-tuning yields limited gains on minimal-pair evaluations, but clearer effects in free generation. Structural alignment produces the strongest improvements in grammaticality, providing a novel, plausible mechanistic account of how this feedback can support grammar learning. Communicative feedback yields more moderate gains. In contrast, semantic contingency and affective feedback do not improve grammaticality, although further analyses suggest that they may support other aspects of language learning beyond grammar. These results suggest that different forms of caregiver feedback make complementary contributions to language learning.
1 Introduction
Caregiver feedback offers children global, noisy signals about whether utterances are communicatively and grammatically successful, but its effects on grammatical learning remain difficult to isolate. This study compares several feedback types in a controlled learning framework.
- Children learn language through social interaction, including caregiver responses to their own utterances.
- Caregiver feedback can act as global positive or negative reinforcement, even though explicit grammatical correction is rare.
- Communicative feedback includes acknowledgments signaling success and clarification requests signaling misunderstanding or failure.
- Semantic contingency measures whether caregiver responses remain responsive to the child’s topic and intended meaning.
- Structural alignment reuses the child’s syntactic frame, while affective feedback varies in positive emotional expression.
- Because feedback types typically co-occur in children’s experience, their causal effects on grammatical learning are difficult to establish experimentally.
- The study compares theoretically motivated feedback types within one controlled learning framework, extending prior work focused mainly on clarification requests.
2 Methods
The study separates linguistic-exposure learning from caregiver-feedback learning by pretraining child-like language models and then applying reward-based fine-tuning. It evaluates both sentence-level grammatical preferences and the grammaticality of generated utterances.
- Models are first pretrained on child-directed language without children’s productions to establish an input-only baseline.
- Models are then fine-tuned with reinforcement learning using reward models trained on caregiver-feedback valence.
- An idealized grammar reward serves as an upper bound on grammatical learning achievable within the framework.
- Minimal-pair evaluations use BLiMP and CHILDES-adapted Zorro to test whether grammatical sentences receive higher probability than minimally different ungrammatical sentences.
- Free generations are evaluated with the CHILDES-adapted ChildCG classifier and a standard grammatical-error-correction metric.
3 Results
Empirical caregiver-feedback rewards produced clearer effects in free generation than in minimal-pair benchmarks. Structural alignment most consistently improved grammaticality, while semantic contingency and affective feedback reduced it.
- The grammar topline improved both minimal-pair benchmark performance and generated-utterance grammaticality above pretraining.
- None of the empirical reward types reliably improved performance across the Zorro and BLiMP minimal-pair benchmarks.
- Structural alignment produced the strongest and most consistent grammaticality improvement across pretraining scales and both ChildCG and GEC metrics.
- Communicative feedback had an overall positive effect, especially under ChildCG, but its gains were smaller than structural alignment’s.
- Semantic contingency and affective feedback generally reduced grammaticality below baseline in generated utterances.
- At 0.1M words, results were noisier and less grammatical; reward-specific effects became more consistent from 1M words onward.
- Communicative feedback improved grammaticality at 1M pretraining words but had a weaker effect at 10M, with clarification requests driving the significant effect.
4 Discussion
The study finds that caregiver feedback types affect grammaticality differently in child-like language models: structural alignment has the clearest benefit, communicative feedback a smaller one, and other feedback types shift language use in different directions. These findings support a controlled account of how naturally occurring feedback may contribute to grammar learning while leaving generalization and multimodal learning unresolved.
- 4 Discussion: Structural alignment systematically improves grammaticality, offering a plausible mechanism whereby caregivers reinforce grammatical syntactic frames by reusing them.The account does not require deliberate teaching because competent caregivers may naturally realign more with well-formed utterances.
- 4 Discussion: Communicative feedback also improves grammaticality, though to a lesser extent than structural alignment.This result is consistent with prior work, especially on clarification requests.
- 4 Discussion: Affective feedback and semantic contingency reduce grammaticality, but supplementary analyses suggest they may support other language properties.Affective feedback increased utterance length, whereas semantic contingency increased content-word use and lexical diversity relative to baseline.
- Minimal-Pair vs. Generation-Based Evaluation: Empirical rewards affected free-generation grammaticality but did not consistently improve minimal-pair benchmark performance, limiting evidence for grammatical generalization.The benchmarks may not cover all grammatical generalizations or the error phenomena common in child production.
- Limitations: The study models only verbal caregiver feedback, although children interpret feedback across linguistic, prosodic, visual, gestural, and facial channels.This multimodal setting could make credit assignment more challenging than in the present framework.
- Limitations: Generality across model architectures remains unestablished because the experiments used a single GPT-2-style architecture.The authors frame the model as a controlled experimental testbed rather than a general-purpose language-model improvement method.
Ethical considerations
The study uses existing child-caregiver corpora and reports aggregate model-level results because the data involve children. The framework is intended for studying language-development hypotheses, not evaluating individuals.
- The study uses existing child-caregiver corpora and does not collect new data.
- Because the corpora involve children, results are reported only in aggregate without identifying individuals or families.
- The modeling framework is designed to study language-development hypotheses, not evaluate individual children or caregivers.
Appendix 1: Supplementary Analyses
The supplementary figures compare baseline, topline, and reward-family performance across pretraining scales. They report benchmark accuracy, generated-utterance grammaticality, and several properties of generated language with seed-based variability.
- Figure 2 compares baseline and topline performance across pretraining scales.The top row reports Zorro and BLiMP minimal-pair accuracy; the bottom row reports generated-utterance grammaticality using ChildCG or GEC-based metrics.
- Figure 3 compares benchmark performance across four empirical reward families.Rows represent Zorro and BLiMP accuracy, while columns represent the reward families; baseline and topline trajectories are repeated for comparison.
- Figure 4 reports sentence length, pooled lexical entropy, function-word density, and non-word rates for generated utterances.Each panel repeats baseline and topline trajectories so reward-specific patterns can be compared across pretraining scales.
Appendix 2: Examples of Model-Generated Utterances
The appendix presents examples generated by reinforcement-learning-fine-tuned models at the 10M-word scale, organized by reward type and classified for grammaticality.
- Table 2 organizes model-generated utterance examples by reward type at the 10M-word pretraining scale.Generation was qualitatively similar at the 1M-word scale.
- Unshaded cells contain utterances classified as grammatical by ChildCG, while shaded cells contain utterances classified as ungrammatical.
Appendix 3: Automatic labeling of rewards
Caregiver-feedback signals were automatically annotated in child-caregiver conversations and used to supervise reward-model training.
- Caregiver feedback signals were automatically annotated in child-caregiver conversations.
- The annotations supervised training of the reward models.Further methodological details are provided in Appendix 4.
Communicative feedback
Communicative feedback was identified from caregiver responses using a fine-tuned DeBERTa-v3-xsmall classifier trained on manually annotated CHILDES data.
- Clarification requests were identified with a DeBERTa-v3-xsmall classifier fine-tuned on manually annotated caregiver responses in CHILDES.
- The clarification-request classifier followed prior work by Nikolaus and Fourtassi (2026).
- The passage also introduces a separate procedure for identifying caregiver acknowledgments using keywords and repetition ratios.
Structural Alignment
Structural alignment measures how much a caregiver response reuses the child utterance’s part-of-speech bigrams.
- A child structure was defined as a part-of-speech bigram, following previous research.
- spaCy was used to assign part-of-speech tags to child utterances and caregiver responses.
- Alignment was the intersection of child and caregiver POS-bigram sets divided by the size of the larger set.
Semantic Contingency
The study operationalizes semantic contingency through embedding similarity and uses shared language-model and reinforcement-learning procedures to train and fine-tune models.
- Semantic Contingency: Semantic contingency was computed as cosine similarity between child-utterance and caregiver-response embeddings, clipped to [0, 1].
- Affective Feedback: Affective properties were annotated for adult responses with roberta-base-go_emotions, using probabilities for supportiveness, approval, and warmth.
- Pretraining: The language models were pretrained on caregiver utterances using a small GPT-2-style causal language model.
- Reward Modeling: Each caregiver-feedback dimension received a reward model mapping a child utterance to the corresponding caregiver-response valence.
- Reward Modeling: All reward models shared a pretrained DeBERTa-v3-xsmall backbone and were fine-tuned with regression and mean squared error loss.
- RL Fine-tuning: After pretraining, models were fine-tuned with PPO by sampling utterances, computing reward-model scores, and updating model weights.