Source-linked AI summary
Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training
Guangqi Li, Yongxin Li
TL;DR
The paper asks how modular task partitions form during LLM training, a process prior work on finished models did not observe. It tracks a Pythia-410M model from scratch with per-step attribution patching and complementary probes across 14 tasks. The map is partly pre-carved, locks in through sharp jumps with gradient-level deprivation that does not reach updates or weights, and remains mechanistically unresolved beyond a construction-level feature account.
Problem
Prior work characterized modular organization in trained models, leaving the formation of task-specific neural populations unobserved.
Method
The study trains Pythia-410M from scratch and applies per-step attribution patching with gradient, update, weight, and loss-decomposition probes across 14 tasks.
Results
The map is pre-carved, the partition locks in through two sharp jumps, and winner neurons receive 2.25→2.73× the loser’s gradient supply without corresponding effective-update or weight asymmetry.
Takeaways & Limitations
The results support a feature-level account in which architecture and input structure shape modularity before learning, while later dynamics amplify the partition.
Takeaways & Limitations
The feature-locking result is established within one determiner–noun agreement construction with three template variants; generality to other constructions and domains is untested.
Abstract
from arXiv · showhide
Large language models exhibit a modular internal organization that mirrors well-studied functional networks of the human brain, but how this organization forms during training is unknown: prior work has characterized finished models, not the formation process. We track formation step by step: we train a Pythia-410M model from scratch (two trajectories, bf16 and fp32) and run attribution patching at every step, alongside probes for gradient norms, effective updates, weight norms, and first-order loss decomposition across 14 tasks in four cognitive domains. Three findings. First, the modular map is pre-carved: before any learning, the dominant task pair already overlaps at ~3.6x the attribution substrate (a task-independent baseline), and its layer-0 concentration is an architecture-level constant on this model family. Second, the partition locks in through two sharp jumps whose amplitudes do not track the learning-rate schedule (the second reaching 20.4 sigma quiet-window / 6.2 sigma global), accompanied by gradient-level relative deprivation--winners receive 2.25->2.73x the loser's gradient supply, 9.5-11.5 standard deviations below a random control--that does not propagate to updates or weights. Third, deviation from the substrate appears only in the domain being learned, consistent with the hypothesis that modularity tracks learning. We close by separating the feature-level account we can defend from the mechanistic questions we cannot, and we pre-register the scale-threshold hypothesis behind our ongoing 2.8B experiments.
1 Introduction
Prior work established modular task organization in trained LLMs but did not observe how task-specific populations form. This paper tracks that process step by step and reports a pre-carved map, sharp lock-in transitions, and gradient-level deprivation without corresponding update or weight asymmetry.
- Prior studies found domain-aligned modularity in trained LLMs, but the formation process remained unobserved at its unfolding resolution.
- Step 0 already showed the dominant pair ≈3.6× above the attribution substrate, with ≈40% layer-0 concentration stable across six corpus variants.
- The partition locked in through two sharp jumps rather than gradual growth, with the second reaching +20.4σ quiet-window / 6.2σ global.
- Winner neurons received 2.25→2.73× the loser’s gradient supply, while effective updates and weight magnitudes showed no corresponding asymmetry.
- Eight quantitative falsifications constrain explanations involving turnover, migration, absolute gradient starvation, irreversibility, weight equilibrium, and three mechanistic factor hypotheses.
2 Related Work
Related work characterizes modularity in trained LLMs and developmental dynamics in smaller or single-task transformer settings. This paper extends the developmental perspective to multi-task modular partitions at per-step resolution and evaluates competing emergence accounts.
- Prior LLM studies localized domain-aligned modular units and ablation effects in trained instruction-tuned models, while other work found emergent structure and recurring components across architectures and scales.
- Developmental transformer studies identified discrete in-context-learning stages and circuit replacement during grokking in small, single-task models.
- This paper extends the developmental lens to multi-task modular partitions in a pretraining-scale model at per-step resolution.
- Interference-avoidance and contingency accounts make opposing predictions, while the reported data indicate input statistics and architecture shape partitions before learning and later dynamics amplify rather than replace them.
3 Method
The study trains Pythia-410M from scratch on 14 tasks across four domains and measures attribution-based task rosters at every step. Substrate calibration, multi-probe measurements, and dual event calibration distinguish task-specific structure from baseline and measurement effects.
- Pythia-410M was trained from scratch in bf16 and fp32 trajectories, with checkpoints, five per-neuron probe arrays, task-conditioned gradients, and attribution tensors recorded every step.
- The 14-task battery spans eight language, two theory-of-mind, and four physical-reasoning tasks built from minimal pairs whose correct continuations flip.
- Attribution rosters contain the top 983 neurons by signed attribution, and pairwise overlap is |A ∩ B|/983 rather than the prior work’s Jaccard metric.
- The attribution substrate is 0.0422, defined as mean step-1 sibling-pair overlap, and deviation is reported as raw overlap / substrate.
- Gradient norms, effective updates, weight norms, activation frequency, second moments, task-conditioned gradients, and first-order loss decompositions provide distinct probe roles against confounds.
- Transitions are detected from first differences and reported using both quiet-window σ and global-MAD calibration, which can produce strongly different significance values.
4 The Map Is Pre-carved
Attribution patching reveals a task-overlap sketch before learning: the regular–irregular pair is elevated above the substrate, while adjective sibling pairs are suppressed. Layer-0 concentration is stable across corpus variants, but the construction-level feature driver remains unresolved.
- The Map Is Pre-carved: At step 0, the regular–irregular pair had 0.153 raw overlap, ≈3.6× the substrate, while adjective sibling pairs had 0.028–0.029 overlap below substrate.
- The Map Is Pre-carved: The step-1 dominant overlap remained similar across trajectories at 0.1404 versus 0.142, with initialization-seed coefficient of variation 0.036.
- The Map Is Pre-carved: Winner tasks placed 39–45% of their rosters in layer 0, invariant across six corpus variants before and after training.
- The Map Is Pre-carved: The adjective variant’s layer-0 attribution was exactly zero on the matched corpus but returned to 0.405 under corpus mismatch, showing corpus dependence.
- The Map Is Pre-carved: Within determiner–noun agreement, structural relation was associated with attribution-pattern switching, but adjacency, intervening-word class, and position could not be separated as the driver.
5 The Partition Locks In via Two Jumps
The partition forms through two sharp overlap jumps while retaining a stable neuronal core amid substantial roster churn. The lock is dynamically maintained rather than a one-shot or purely additive transition.
- 5.1 A two-jump structure, not gradual growth: +0.154 overlap at step 79→80 marks the largest jump, reaching 20.4σ quiet-window / 6.2σ global-MAD significance.A first jump occurred at step 18→19, followed by consolidation and a second transition that carried overlap to 0.53–0.61.
- 5.2 The lock is a high-churn reorganization, not stable amplification: +71% (run2) / +121% (run1) shared-set growth coincided with only 60%/44% old-member retention, ruling out pure expansion or replacement.Both runs used the same frozen two-threshold criterion.
- 5.2 The lock is a high-churn reorganization, not stable amplification: 0/117 per-step adjudication cases showed turnover or migration across the full window, despite substantial neuron-level flux.The sole turnover case in the entire grid was the adjective task at step 79.
- 5.2 The lock is a high-churn reorganization, not stable amplification: 165 neurons form a persistent skeleton, while 1,124 neurons repeatedly enter and exit the winner shared set.The skeleton constitutes at least 90% of the shared set from step 24 onward, but membership remains fluid.
- 5.2 The lock is a high-churn reorganization, not stable amplification: The run1 conflict–cleaning cycle is not general: run2 dips were one-third to one-half as deep, never significantly below control, and lacked large pruning.The authors report this cycle as a property of run1 rather than a universal law.
6 Relative Deprivation That Does Not Propagate, and a Coupling to Learning
Winner neurons receive substantially more gradient supply than the loser, but this asymmetry remains confined to gradients rather than effective updates or weight norms. Domain-level deviation appears where loss is being learned, linking modularity to learning without establishing deprivation as causal.
- 6.1 Gradient-level deprivation: 2.25→2.73× winner-to-loser gradient supply remained 9.5–11.5σ below a random same-size control.The loser/winner gradient-norm ratio fell from 0.44 to 0.37 across training.
- 6.1 Gradient-level deprivation: ≈1 effective-update and weight-norm ratios show that gradient deprivation does not propagate to Adam-normalized updates or parameter magnitudes.Weight norms drifted only 0.40–0.82% over the trajectory.
- 6.1 Gradient-level deprivation: The authors do not claim starvation as the partition mechanism because deprivation is associated with exclusion but its causal direction is unsettled.The loser remains parameter-healthy and later exhibits weak learning.
- 6.2 Coupling to learning: −0.29 →−1.77 language loss components contrasted with physics and theory-of-mind remaining structurally flat at ±0.05.Both lock-in jumps occurred exclusively in the language domain.
- 6.2 Coupling to learning: ≈1.02 physics within-domain/cross-domain overlap stayed inside [0.8, 1.25] for 81/85 steps, indicating no comparable physics-specific deviation.The result accompanies the domain-selective first-order loss pattern.
7 Discussion
The discussion separates a supported feature-level and dynamics-level account from unresolved mechanism-level questions. The results favor pre-existing structural alignment and amplification over interference-driven origin, while constraining future explanations with falsifications.
- Three levels of answer: At the feature level, the step-0 map reflects structural separation, output word-form, and an architecture-level layer-0 concentration band.The paper explicitly limits this account to statistical alignment rather than circuits at initialization.
- Three levels of answer: At the mechanism level, the source of attribution alignment in random networks, the ≈40% band, and the substrate floor remains unanswered.The authors avoid mechanistic vocabulary that would imply an unsupported explanation.
- Relation to existing accounts: Partitions existing before learning rule out interference avoidance as the origin, while reproducible topology with variable roster identity and timing supports only a weak contingency account.The evidence distinguishes stable partition structure from contingent membership and event timing.
- Falsifications as contributions: 0/117 turnover or migration cases, spontaneous recovery, and r = −0.27 between supply and norms constrain turnover, irreversibility, and weight-equilibrium explanations.The paper reports eight quantitative falsifications intended to narrow future mechanism theories.
8 Limitations
The paper’s limitations constrain the generality of its feature and dynamics claims, leave trajectory differences and deprivation mechanisms unresolved, and make conclusions protocol- and scale-specific.
- Feature-locking evidence comes from one determiner–noun agreement construction with three template variants, leaving other constructions and domains untested.
- Dynamics claims rely on two trajectories of a single 410M model, without cross-seed training replication.
- The bf16 and fp32 trajectories differ in unresolved ways, including conflict-cycle depth, event significance regime, and roster fluidity.
- Relative deprivation is correlational because supply equalization has not established its causal role.
- At 410M, the physics domain is not learned, so scale-threshold conclusions remain preregistered rather than tested.
- The two localization calibers are non-interchangeable, making conclusions caliber-specific by protocol.
Ethics Statement
The supplied passages describe the paper’s measurement discipline and scope rather than a conventional ethics statement. They emphasize calibrated, frozen protocols and explicitly bounded mechanistic conclusions.
- The study uses substrate-calibrated deviation, dual-calibration significance, and multi-probe cross-checks to interpret attribution dynamics.
- A domain-blind substrate was falsified because all 12 tested units were domain-specific, motivating domain-specific calibration.
- 15.26× random equals approximately 3.6× substrate deviation, so random multiples are not used as overlap effect sizes.
- Signed and absolute-value localization calibers are non-interchangeable, and the main text uses the signed caliber frozen in a written protocol.
- Gradient-roster conclusions retain explicit caliber labels because |grad norm task| and |grad task| orderings are not equivalent.
- The paper reports eight quantitative falsifications to constrain candidate formation mechanisms and future theory.
Appendix C: Anomaly Registry (Paper-Relevant)
The anomaly registry documents unresolved trajectory, caliber, measurement, and data-driven events while preserving the paper’s main calibrated comparisons and exclusions.
- Anomaly Registry: The adjective task’s layer-0 attribution is absent across all 98,304 neurons, but a mismatch corpus restores 405/395 members, making the phenomenon corpus-dependent.
- Trajectory differences: Cross-trajectory random-control behavior reverses between runs, leaving the trajectory-identity question unresolved because seed, learning rate, and batch identity are unconfirmed.
- Anomaly Registry: The V spike pair at steps 114→116 reaches −0.151/+0.149 and 19.8σ quiet-window, but remains unattributed.
- Caliber disputes: Physics entry conclusions remain caliber-conditional because the protocol and archive disagree over gradient-roster definition, whose orderings are not equivalent.
- Probe anomalies: At step 108, gradient-cosine 0.93 holds only for the |grad task| amplitude-profile max-pair caliber, while signed mean-pair cosine is near zero.
- Probe anomalies: The step-108 retention discrepancy is 49.34% versus 57.0%, reflecting a definition/caliber issue under review.
- Anomaly Registry: The anomaly registry records a dual-caliber identity split for the adjective task: attribution membership is fluid while the loss-decomposition working set is stable.
- Anomaly Registry: The claim that the module is pinned to layers 0–7 is corrected to a mass characterization because the shared set spans layers 3–19 and reaches layer 8 at step 1.
Appendix F: 22-Task Orthogonal Family (Design Frozen, Not Yet Run)
Appendix F freezes a 22-task orthogonal family and a response-surface regression to disentangle candidate feature axes underlying the construction-level separation. It also specifies scale-threshold signatures and a 2.8B protocol, while reporting that physics modularity is absent at 410M.
- Design rationale: Three sibling pairs cannot disentangle template similarity, output word-form, and dependency span, motivating the frozen 22-task orthogonal family.The design is hash-bound but not yet executed; frozen fp32 checkpoints permit attribution back-filling at steps 0/1/25/50/85.
- Design rationale: Five axes vary template similarity, output word-form, dependency span, answer vocabulary, and semantic pairs in a frozen 22-row feature×task matrix.The matrix is transcribed in the design document.
- Design rationale: ≈85% of roster members are replaced by answer-token swaps in synthetic corpora, but real sibling-pair suppression is only 3.7–4.5×, motivating real-family tests of word-form effects.Synthetic answer-token variants reach overlaps 0.910/0.887/0.897, approximately 21× substrate, whereas the base answer pool reaches 0.135–0.229.
- Planned analysis: The pre-registered response-surface model regresses step-0 attribution-overlap statistics on five axes, with axis-3 span versus axis-2 word-form coefficients testing the unresolved contradiction.The regression has coefficients β1–β6 and has not yet been run.
- Scale protocol: The 2.8B pre-registration tests whether module formation requires learning plus deviation that follows learning, using S4/S5/S6 measurements and four frozen physics branches.The protocol includes an early 50-step self-training run and candidate hypotheses involving ICL thresholds, domain-level deprivation, and capacity or landscape effects.
- Scale protocol: At 410M, physics loss is near-flat and its within/cross overlap ratio is 1.0199 median, with 81/85 steps inside [0.8, 1.25], while a steps-19–23 burst decays immediately.The reported scale-threshold lower bound is (1.4B, 2.8B], based on physics modularity being absent at 410M and 1.4B and present at 2.8B according to official checkpoints.