Source-linked AI summary
Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans
Jinglei Ren, Yuyue Wang
TL;DR
It remains unclear whether language models reactivate prior representations or re-evaluate repeated words, and whether post-training changes this behavior. Using repetition priming across two tasks, 15 models, and matched human experiments, the paper finds a qualitative split: base models show automatic facilitation, whereas instruct models show context-dependent processing that can become interference at larger scales.
Problem
It remains unclear whether language models reactivate prior representations or re-evaluate repeated words, and whether post-training changes this default processing behavior.
Method
The study applies repetition priming to 15 models across five families in semantic categorization and cloze completion, with matched human experiments using identical stimuli.
Results
Base models show immediate, lag-stable facilitation, while instruct models show weaker lag-decaying facilitation that reverses to interference at larger scales across tasks.
Takeaways & Limitations
Repetition priming reveals a qualitative processing shift associated with post-training, while humans show a hybrid profile that neither model type captures.
Takeaways & Limitations
Raw log-probabilities may reflect sentence-level entropy differences, although within-item baselines and rank-normalized analyses leave the results qualitatively unchanged.
Abstract
from arXiv · showhide
Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior. We apply repetition priming (Shiffrin and Schneider, 1977) to 15 models across five model families (1.5B-14B parameters) in two tasks, semantic categorization and cloze completion, with matched human experiments using identical stimuli. We find that base models exhibit automatic processing: they show immediate facilitation that remains stable across lags, partially survives context removal, and correlates with attention to prior occurrences. Instruct models exhibit controlled processing: their facilitation decays with lag, collapses without expected context, and reverses to interference at larger scales. Within the Qwen 2.5 family, this dissociation increases monotonically with model scale, suggesting that post-training progressively alters repetition processing. Humans show a hybrid profile, with lag-sensitive facilitation resembling instruct models but without interference, suggesting that neither model type fully captures human cognition. Our findings reveal a qualitative shift in how language models process repeated information after post-training and provide mechanistic evidence for the divergence between model behaviors.
1 Introduction
Using repetition priming, the study identifies a qualitative post-training shift in default processing: base models show automatic-like facilitation, whereas instruct models show controlled, context-sensitive processing. This dissociation appears across two tasks and is supported by mechanistic probes of context, label transfer, and attention.
- Motivation and framework: Repetition priming probes whether repeated words reactivate prior representations or are re-evaluated afresh.The framework contrasts automatic processing, which is fast and interference-robust, with controlled processing, which is context-dependent and disruption-sensitive.
- Experimental design: The study tests 15 models across 5 families spanning 1.5B–14B parameters in semantic categorization and cloze completion.Matched base/instruct variants enable controlled within-family comparisons while holding architecture and pretraining family as constant as described in the passage.
- Core findings: Base models show strong, immediate facilitation that remains stable across lags and robust to intervening material.This profile is consistent across both tasks.
- Core findings: Instruct models show weaker facilitation that decays with lag and reverses to interference at larger scales.For instruct models, a repeated word that initially helps the model can later hurt it.
- Mechanistic evidence: Mechanistic probes find residual facilitation after removing prior response context in base models but interference in instruct models, with attention ablation reducing priming only in base models.Label remapping transfers partially in base models and not at all in instruct models.
2 Related Work
Prior work distinguishes task-level in-context learning and related-word or structural priming from same-word repetition priming. Studies also indicate that instruction tuning can produce deep representational changes, motivating comparisons across model families and scales.
- In-context learning: In-context learning adapts model behavior from prompt examples without parameter updates, whereas this study examines item-level repetition across intervening material.Mechanistic ICL work identifies attention-based retrieval, including induction heads, and characterizes ICL as implicit optimization.
- Priming paradigms: Semantic priming concerns related words and structural priming concerns reused syntactic constructions, while repetition priming targets the same word.Natural and ICL-induced repetition in generation have been reported to rely on distinct mechanisms involving confidence, attention allocation, and perturbation sensitivity.
- Study scope: The study tests 15 models across 5 families, with Qwen 2.5 spanning four parameter scales to examine how model size interacts with instruction tuning.The model set is designed to support analysis of instruction tuning and scale in information processing.
- Instruction tuning: Instruction tuning can reshape model behavior through changes to self-attention heads and feed-forward representations, suggesting deep representational changes rather than superficial formatting shifts.Comparisons of base and instruct variants are described as showing systematic behavioral differences, but the supplied passage is truncated before the specific comparison result.
3 Experiment 1: Semantic Categorization
Experiment 1 used repeated semantic categorization to compare 15 base and instruct models across five families and matched human participants. Base models showed strong, lag-invariant facilitation, whereas instruct models showed weaker, lag-sensitive priming that could reverse to interference; humans showed lag-sensitive facilitation without interference.
- Experiment design: The experiment tested 15 models spanning five families and five parameter scales, including matched base and instruct variants and Qwen 2.5 models from 1.5B to 14B.Each target appeared four times at lags of 1, 3, 7, or 15 intervening trials, with context reset before each word.
- Model results: All seven base models showed strong positive final-exposure priming, with E4−E1 values ranging from +5.04 to +6.89.All effects were significant at p < .001.
- Model results: All seven base models showed positive exposure effects without lag effects, while six of eight instruct models showed negative lag effects and five survived FDR correction.In Qwen2.5-14B-Instruct, facilitation reversed to interference at larger lags.
- Model results: Within Qwen 2.5, base priming remained stable across scale, whereas instruct priming declined monotonically and reversed to interference at 14B.The passage cautions that this matched-family trend does not establish size alone as the explanation for cross-family variation.
- Human comparison: Humans showed robust facilitation: mean RT decreased from 648 ms at E1 to 539 ms at E4, while accuracy increased from 92.3% to 97.1%.RT priming decayed with lag (β = −1.72), whereas accuracy priming was not significantly lag-sensitive (β = −0.08, p = .16).
- Human comparison: Human lag sensitivity resembled instruct models rather than base models, but humans remained uniformly facilitative and showed a low-frequency RT advantage absent in LLMs.The low-frequency advantage was 131 ms versus 87 ms (p = .003).
4 Experiment 2: Cloze Completion
Experiment 2 tested repetition priming in open-ended cloze completion using weakly constraining contexts, revealing strong, lag-insensitive facilitation in base models but weaker, lag-sensitive effects in instruct models. Humans showed gradual, lag-sensitive facilitation that remained above baseline even at the longest lag.
- Design: The experiment used 21 target words in four low-cloze sentence frames each, ensuring sentence context weakly constrained the target and isolating prior in-context exposure.Target words had cloze probabilities of 0.05–0.15 and zero contamination; each appeared at exposures E0–E3 with lags of 1, 3, or 7 filler trials.
- Model results: Base models showed strong positive priming, with E3−E0 values ranging from +2.11 to +4.12, whereas instruct models ranged from +0.61 to +1.85.All seven base models had p < .001, and the base/instruct gap appeared across all model families.
- Model results: All base models showed significant exposure effects without lag sensitivity, while all eight instruct models showed negative lag effects and six remained significant after FDR correction.ModelType×Lag interactions were confirmed in Experiment 2 (β = −0.13, SE = 0.03, t = −4.52, p < .001).
- Model results: Within Qwen 2.5, base-model priming improved with scale while instruct-model priming declined, replicating the Experiment 1 scaling pattern.The scaling trend was observed within the Qwen 2.5 family.
- Human results: Humans’ target production increased from 11.2% at baseline to 28.4% after one exposure and 41.7% after three exposures, while response times fell from 4.82 s at E1 to 3.41 s at E3.Both production rate and response time showed significant lag-dependent decay, but production remained above baseline at the longest lag: 19.8% vs. 11.2%.
5 Mechanistic Evidence
Mechanistic analyses show that base and instruct models both retrieve prior occurrences, but diverge in how retrieval shapes behavior. Base-model facilitation depends causally on prior-occurrence attention, whereas instruct-model outcomes are decoupled from retrieval and increasingly suppress it with scale.
- Context ablation: Base models retained positive priming without response history, whereas instruct models shifted to negative priming, revealing weakened facilitation versus context-dependent interference.Base models averaged +0.70, while instruct models averaged −2.36; Qwen2.5-14B-Instruct reversed from positive priming at lag 1 to a mean of −0.61.
- Label remapping: 3 of 7 base models showed significant facilitation after label remapping, while no instruct model showed significant transfer, indicating word-level versus stimulus-specific priming.The remaining four base models showed non-significant positive trends after reversed labels.
- Attention retrieval: Both model types attended above baseline to prior occurrences, but base models showed consistently elevated attention while instruct models were more variable.Attention ratios were lower in Experiment 2, where longer cloze contexts distributed attention more broadly.
- Attention–behavior relationship: Base-model attention to prior occurrences correlated positively with priming, whereas instruct-model correlations were weak and non-significant, with |r| < 0.12.This indicates that retrieval directly drives base-model facilitation but is decoupled from instruct-model behavioral outcomes.
- Scaling trend: Within Qwen2.5 instruct models, attention ratios declined monotonically with scale, and the largest models allocated less attention to prior occurrences than chance.This suggests that post-training may increasingly suppress attention-based retrieval at sufficient scale.
- Causal ablation: LLaMA3 base-model priming fell from +6.89 to +2.62 after top-head ablation, while instruct-model priming changed only from +2.05 to +1.97.Qwen2.5 base models similarly fell from +5.52 to +1.84, whereas instruct models changed from +2.31 to +2.14; random-head controls largely preserved base priming.
6 Discussion and Conclusion
The discussion frames repetition priming as a dissociation between automatic processing in base models and controlled, context-dependent processing in instruct models. Humans combine instruct-like lag sensitivity with base-like facilitation, while the findings have implications for systems that reuse contextual information.
- Automatic vs. controlled processing: Base models show automatic processing, with lag-invariant facilitation that partly survives context removal and label remapping.The results are interpreted within the automatic-versus-controlled processing framework of Shiffrin and Schneider (1977).
- Automatic vs. controlled processing: Instruct models show controlled processing, re-evaluating repeated words in a context-dependent manner that decays with lag and collapses without expected context.Correlational attention results and causal head ablations converge on a role for prior-occurrence retrieval.
- Human comparison: Human processing is lag-sensitive like instruct models but uniformly facilitative like base models.This hybrid profile is discussed with reference to Bowers (2000) and Tenpenny (1995).
- Practical implications: Repetition effects in RAG, few-shot prompting, and dialogue are predicted to depend on model type and recency when contextual information is reused.The relevant repeated units include chunks, examples, and entities.
Ethics Statement
The human experiments were IRB-approved, conducted with informed consent and compensation, and involved standard psycholinguistic tasks without deception, sensitive content, or foreseeable risks beyond everyday computer use.
- The human experiments were approved by Yale University’s Institutional Review Board, and all participants provided written informed consent.
- Participants were compensated at a rate of $15/hour.
- The study used standard word categorization and sentence completion tasks without deceptive procedures, sensitive content, or foreseeable risks beyond everyday computer use.
LLM Usage Disclosure · A Extended Related Work
The paper discloses author-verified use of an LLM writing assistant for language editing and camera-ready organization. Its related work situates the study within dual-process LLM frameworks, LLM–human cognitive comparisons, instruction-tuning mechanisms, neuroscience of repetition suppression, and distinct repetition mechanisms.
- LLM Usage Disclosure: An LLM-based writing assistant supported language editing and camera-ready organization, while authors reviewed and verified all scientific claims, analyses, and final text.
- A Extended Related Work: Dual-process frameworks distinguish fast heuristic System 1 from slow deliberate System 2 processing in LLMs.Prior work links LLM behavior to prompting strategy and frames reasoning-model progress as a shift toward System 2 capabilities.
- A Extended Related Work: Research increasingly evaluates LLMs as models of human cognition, reporting both striking correspondences and systematic divergences.Studies benchmark humanlikeness across linguistic dimensions and examine behavioral patterns such as semantic priming and garden-path effects.
- A Extended Related Work: Instruction tuning reshapes model internals, including self-attention relationships and feed-forward knowledge directions.Mechanistic interpretability work suggests fine-tuning surfaces pre-existing capabilities rather than constructing entirely new circuits.
- A Extended Related Work: Repeated stimuli typically produce repetition suppression in neural responses alongside behavioral facilitation.This neuroscience dissociation parallels the paper’s finding that instruct models can reduce attention to prior occurrences while retaining some behavioral facilitation.
- A Extended Related Work: Human and language-model predictions diverge for repeating passages, with LM patterns linked to middle-layer attention heads and recency.Other work distinguishes natural repetition from in-context-learning-induced repetition through confidence, specialized heads, and perturbation sensitivity.
B Experimental Design Details … E Model Specifications
The experiments isolated within-word repetition by manipulating lag and context, using semantic categorization and cloze completion with distinct stimulus sets. Prompt controls and standardized inference settings tested whether model behavior depended on wording or implementation choices.
- B Experimental Design Details: Lags were tested separately with fresh context: 1, 3, 7, and 15 in Experiment 1, versus 1, 3, and 7 plus an isolated baseline in Experiment 2.Each lag condition was run as a separate session.
- B Experimental Design Details: Each target word appeared in an independent context block, with context reset before first exposure to isolate within-word repetition.Experiment 1 used single-word classification with 45 filler trials; Experiment 2 used sentence completion.
- C Stimuli: Experiment 1 included 182 McRae-norm target words, evenly divided between 91 living and 91 nonliving items.The stimuli were selected from McRae feature norms.
- D Prompt Formats: Three semantically equivalent Experiment 1 templates produced LLaMA3-Base lag coefficients from −0.04 to +0.01 and LLaMA3-Instruct coefficients from −0.32 to −0.27.Base coefficients were all n.s.; Instruct coefficients were all p < .001.
- D Prompt Formats: The Prompt×ModelType×Lag interaction was nonsignificant (F = 1.12, p = .33), indicating wording differences did not explain the model-type dissociation.The templates varied instruction wording and answer cues while preserving response options.
- D Prompt Formats: Additional controls replaced responses with “...” placeholders and tested abstract A/B labels across remapped mappings, warmup trials, and target-versus-control words.The label-remapping procedure used 20 words × 3 exposures, a mapping change, 8 warmup trials, and 20 TARGET versus 20 CONTROL words.
- E Model Specifications: All experiments used HuggingFace Transformers v4.36.0, bfloat16 precision, eager attention, greedy decoding, a 100-trial context limit, and NVIDIA A100 80GB GPUs.These settings standardized model inference across experiments.
F Statistical Analysis Details · G Human–Model Comparison Details
The statistical analysis defines exposure-specific and mean priming effects, tests ModelType, Exposure, Lag, and interactions with mixed-effects models, and reports large between-type contrasts. Human–model comparisons are profile-based: they assess direction and lag dependence rather than comparing human and model measures in absolute magnitude.
- F Statistical Analysis Details: Priming is computed as MarginEi − MarginE1 for E2−E1, E3−E1, E4−E1, and their mean across E2–E4.Values are averaged across all target words and lag conditions.
- F Statistical Analysis Details: The ModelType×Lag interaction was significant in Experiment 1 (β = −0.26, SE = 0.04, t = −6.50, p < .001) and Experiment 2 (β = −0.13, SE = 0.03, t = −4.52, p < .001).The omnibus mixed-effects model included ModelType, Exposure, Lag, and their interactions as fixed effects, with random intercepts for Model and Word.
- F Statistical Analysis Details: False-discovery-rate correction preserved negative lag effects for 5/8 instruct models in Experiment 1.The supplied passage truncates the corresponding Experiment 2 statement.
- F Statistical Analysis Details: Experiment 1 mean priming d = 2.18, Experiment 1 final priming d = 1.82, Experiment 2 mean priming d = 2.90, and the missing-context contrast d = 2.93.These standardized between-type contrasts were large.
- F Statistical Analysis Details: Attention analysis identifies prior presentations, sums final-token attention to them, averages across heads per layer, and computes the ratio to uniform-distribution expectation.Targeted ablation ranked heads by the association between attention to earlier target occurrences and item-level priming.
- F Statistical Analysis Details: Ablating the top 10 retrieval-associated heads reduced base-model priming, whereas instruct models and random-head controls changed little.The procedure targeted LLaMA3-8B and Qwen2.5-7B base/instruct pairs, with random sets of 10 heads as controls.
- G Human–Model Comparison Details: Human and model measures are not comparable in absolute magnitude; human RT, accuracy, and production instead provide a qualitative cognitive reference.Profile-level comparisons assess direction and lag dependence rather than metric-level magnitudes.
- G Human–Model Comparison Details: Human summaries compare E1 with E4 in Experiment 1, no-study baseline with three exposures in cloze production, and E1 with E3 among correctly produced targets in RT.These comparisons define the human performance profiles used in the human–model analysis.
H Additional Results
Additional analyses show that repetition priming varies with exposure and lag across both experiments. Ablations and label remapping further indicate divergent priming patterns and representational facilitation.
- Exposure effects: Priming by exposure was evaluated in both experiments relative to their respective baseline measures.Experiment 1 used change in log-probability margin from the E1 baseline; Experiment 2 used change in target log-probability from the E0 baseline, averaged across lag conditions.
- Exposure effects: Mixed-effects analyses modeled exposure and lag effects in Experiments 1 and 2.Exposure β represented change per additional exposure, while Lag β represented change per unit lag increase.
- Ablation analysis: With responses replaced by placeholder tokens, base models showed weak positive priming whereas instruct models showed strong negative effects.This was assessed in a no-response ablation by lag in Experiment 1.
- Ablation analysis: Qwen2.5-14B-Inst showed +2.49 priming at lag 1 before reversing to negative priming at longer lags.The reversal occurred under the no-response ablation in Experiment 1.
- Label remapping: Under remapped labels, LLaMA3-Base showed higher TARGET than CONTROL margins, indicating representational facilitation.The accompanying learning curve showed a rapid margin increase during Phase 1.