Source-linked AI summary

Where Grokking Happens: Distributed Utility and Fourier Recoding Without a Module Switch

Dekun Yang

arXiv:2609.17571v1cs.LGcs.AI

TL;DR

The functional locus of the memory-to-generalization transition remains unclear because endpoint localization and output-space Fourier progress do not allocate longitudinal change among interacting components. Transition Games compare behavior-anchored checkpoints using exact activation-replacement Shapley games, matched non-generalizing controls, Fourier subplayer refinement, donor replacement, and exact downstream path restoration. Grokking shows distributed, unequal utility gains with a prospective block-0 attention bias; block-1 MLP mediates more downstream effect than the other three tested paths combined.

  • Problem

    The functional locus of the memory-to-generalization transition remains unclear because endpoint localization and output-space Fourier progress do not allocate longitudinal change among interacting components.

  • Method

    Transition Games compare behavior-anchored checkpoints using exact activation-replacement Shapley games, matched non-generalizing controls, Fourier subplayer refinement, donor replacement, and exact downstream path restoration.

  • Results

    Grokking shows distributed, unequal utility gains with a prospective block-0 attention bias; block-1 MLP mediates more downstream effect than the other three tested paths combined.

  • Takeaways & Limitations

    Within these small Transformers, grokking is spectral recoding of an existing distributed circuit rather than a winner-take-all module switch or universal architecture law.

  • Takeaways & Limitations

    The evidence is limited to small tied-embedding Transformers and hybrid activation interventions whose allocations do not identify a literal operator or unique route.

Abstract

from arXiv · show

Where in a Transformer is the change from memorization to generalization functionally expressed? We introduce Transition Games--behavior-aligned exact activation games with paired non-generalizing controls--and find distributed utility gain with a prospective block-0 attention bias; selected degree-two modes account for 67--92% of its addition contrast across replacement games, and a disjoint exact path study confirms that block-1 MLP mediates more of their effect than all other tested downstream paths in 12/12 pairs. The sharper "MLP memorizes, attention generalizes" prediction instead reverses (-.331 bits/example at the memory anchor; 0/12 in the predicted direction), while routing onset, global rank collapse, and a prime-invariant architecture ridge also fail, identifying grokking here as spectral recoding of an existing distributed circuit rather than a module switch.

1 INTRODUCTION

The paper asks where grokking’s memory-to-generalization transition is functionally expressed and measures that transition directly with paired exact activation games. It finds unequal gains across a distributed circuit, with spectral recoding and downstream block-1 MLP mediation rather than a module switch.

  • Transition Games: Transition Games compare the last memorizing OOD-poor checkpoint with the first high-OOD checkpoint, subtracting matched non-generalizing control changes from exact Shapley allocations.The estimand concerns the declared activation-replacement game, not a unique parameter-level storage address.
  • Structural localization: All three residual paths gain OOD utility, but block-0 attention gains relative to neighboring MLP across prospective addition and division cohorts.The attempted MLP-memory signature reverses at memorization, the strict memory anchor, and generalization; routing onset and global rank collapse also fail.
  • Spectral recoding: Selected embedding-supported degree-two Fourier modes account for most of block-0 attention’s controlled addition gain, become more decodable on held-out inputs, and transfer to division.The evidence persists across frequency selection and deterministic marginal-preserving donor replacement, although donor–recipient hybrids are not fully on-manifold.
  • Path mediation: Exact restoration on disjoint seeds shows block-1 MLP mediates more selected-mode effect than the other three tested downstream paths combined.The frozen block-0 MLP-versus-direct comparison was null, while the required full allocation generated and confirmed the final-block MLP hypothesis.
  • Architectural scope: A prime-by-width-by-head factorial supports a predeclared head-dimension diagonal at modulus 43 but not 53, rejecting a universal architecture ridge without identifying its causal moderator.Matched training density and leave-one-out analyses rule out two simple explanations, while transition delay correlates with prime but not with the effect within either prime.

2 RELATED WORK

Related work distinguishes endpoint localization, output-space Fourier progress, and other longitudinal or circuit methods from the paper’s behavior-anchored paired exact activation game. The paper positions its contribution as a narrower allocation and mediation estimand rather than a complete theory of grokking.

  • Fourier and grokking: Output-space Fourier measures establish algorithmic progress but cannot determine which interacting component changes its functional contribution across grokking.The paper’s longitudinal game supplies that allocation while avoiding a claim to a complete theory of grokking.
  • Longitudinal methods: Contemporaneous longitudinal methods attribute components through gradient-path kernels or track Fourier-feature alignment, whereas Transition Games estimate behavior-anchored paired exact activation changes and downstream mediation.The comparison is methodological: these approaches offer complementary views rather than estimating the same quantity.
  • Architectural accounts: The paper’s factorial crosses width and head count in standard attention–MLP blocks at nearly fixed training density, so its prime-specific result rejects extrapolating one head-dimension optimum across tasks.It neither replicates nor refutes the distinct attention-only boundary described in prior work.
  • Localization and editing: Endpoint importance, recall, editing, or causal-tracing results need not identify the component whose role changes during generalization or the best layer for editing.This motivates distinguishing allocation in a stated intervention game from endpoint storage and editability.
  • Circuit and intervention methods: Activation patching, causal scrubbing, ACDC, and attribution patching provide circuit-discovery tools, while exact Shapley and Owen games define allocations through explicit players, replacements, utilities, and sampling schemes.The paper emphasizes that intervention-game design exposes rather than erases dependence on these choices.

3 EXPERIMENTAL DESIGN

The study operationalizes grokking as a controlled change in activation-game utility between matched memory and generalizing checkpoints, using exact coalition and path analyses across tasks, replacements, and architectures.

  • 3.2 TRANSITION GAMES: EXACT LONGITUDINAL ACTIVATION GAMES: Transition Games compare the last memory-dominant checkpoint with the first persistent-generalization checkpoint, subtracting matched non-generalizing control changes from exact activation allocations.The estimand targets the declared activation-replacement game rather than claiming a unique parameter-level locus.
  • 3.2 TRANSITION GAMES: EXACT LONGITUDINAL ACTIVATION GAMES: Exact Shapley games allocate clipped predictive compression across embedding, attention, and MLP activations, with complete-coalition reconstruction and efficiency audits.Inactive players are replaced by training-set means or zeros, and utility is measured in bits/example.
  • 3.3 STRUCTURAL, HANDOFF, AND FACTORIAL TESTS: The factorial study crosses widths 128/256 with four/eight heads at primes 43 and 53, requiring the predeclared diagonal contrast to have the predicted sign at both primes.The cross-prime claim therefore tests architectural universality rather than merely reporting a single-prime effect.
  • 3.3 STRUCTURAL, HANDOFF, AND FACTORIAL TESTS: Selected degree-two Fourier coordinates refine block-0 attention through exact Owen values, with held-out probes, selective removal, and cross-operation transfer to division in multiplicative-group coordinates.The addition analysis nests Fourier children inside the layer game, while division uses ratio modes after mapping operands through a primitive root.
  • 3.5 PATH RESTORATION: A four-site restoration game tests whether the selected block-0 attention effect follows direct, block-0 MLP, block-1 attention, or block-1 MLP paths under disjoint prospective path outcomes.The mandatory full allocation generates the block-1 MLP hypothesis before the frozen two-test Holm family is opened.
  • 3.7 INFERENCE AND AUDITS: The analysis uses seed/split pairs, frozen test families, one-sided exact sign-flip tests with Holm correction, and deterministic replay across primary and sensitivity games.Marginal-preserving donor replacement preserves each target activation’s empirical marginal while breaking within-example dependence, though it still creates hybrid states.

4 RESULTS

Grokking changed all top-level paths, with a reproducible relative block-0 attention advantage rather than an MLP-to-attention handoff. Selected degree-two modes captured most of the addition contrast, and their effect was mediated chiefly through block-1 MLP.

  • 4.1 THE TRANSITION IS DISTRIBUTED, WITH A RELATIVE ATTENTION BIAS: Attention gained .1874 more than MLP ([.1322, .2425]; 12/12; two-sided exact p = .000488), while every top-level path gained OOD utility in all 12 pairs.Embedding gained .1228 more than MLP, but attention and embedding were not significantly separated.
  • 4.2 THE PROPOSED MODULE HANDOFF REVERSES: At the strict memory anchor, MLP-minus-attention training utility was −.3314 bits/example ([−.4648, −.1980]; 0/12 in the predicted direction), reversing the proposed handoff.The reversal persisted under zero replacement, unclipped log-likelihood advantage, and accuracy allocation.
  • 4.3 FRESH COHORTS CONFIRM THE RELATIVE BLOCK-0 LOCUS: The prospective block-0 attention-minus-MLP gain was .1632 bits/OOD example for addition and .1591 for division, with none of 24 controls generalizing.The addition cohort passed its behavioral gate in 12/12 runs and division in 11/12.
  • 4.4 EXACT SUBSPACE GAMES IDENTIFY SPECTRAL RECODING: Selected degree-two modes contributed 78.8% of the aggregate addition contrast under mean replacement and 91.5% under zero replacement, with a positive residual.Donor replacement independently yielded 67.1% selected/group shares, so the account is not exclusive to five frequencies.
  • 4.5 SELECTED-MODE EFFECTS FLOW MAINLY THROUGH THE FINAL MLP: On disjoint seeds, block-1 MLP gained 2.6440 bits and exceeded the other tested paths combined by 2.5167 in 12/12 pairs.The result localizes mediation from block-0 attention through the final MLP, not a literal degree-two operator.
  • 4.6 ARCHITECTURAL UNIVERSALITY FAILS: The cross-prime architecture ridge was rejected: the diagonal was .2115 bits/example at p = 43 but .0191 at p = 53.The moderator remained unidentified, and no architecture recommendation was made.

5 DISCUSSION

The discussion interprets grokking as task-aligned spectral recoding within a distributed circuit, not a winner-take-all module switch. It also limits the claim to declared games, tested architectures, and partial mechanistic localization.

  • 5 DISCUSSION: All coarse paths gain OOD utility, attention leads relatively, and selected-mode evidence traces chiefly through block-1 MLP rather than defining a generalization module.The final-MLP dominance is conditional on one block-0 spectral intervention.
  • 5 DISCUSSION: TRANSITION GAMES distinguish coordinated internal allocation from output-space Fourier progress, and the unseen final-MLP prediction was confirmed after the base path target remained unresolved.This makes the negative handoff result informative rather than merely eliminative.
  • 5 DISCUSSION: Selected modes carry 67–92% of the addition contrast across replacement games, while residuals remain and division favors broad same-frequency context over pure ratio lines.Path restoration identifies block-0-attention-to-final-MLP mediation, not a literal multiplication instruction or unique decomposition.
  • 5 DISCUSSION: Alternative utilities agree on the central addition, handoff, selected-subspace, donor, and path findings but expose division and factorial boundaries.Those disagreements preclude selecting only the favorable metric.

6 LIMITATIONS

The evidence is bounded by small tied-embedding Transformers, limited operation and architecture coverage, and intervention-specific interpretability assumptions. These constraints leave broader transfer, moderator identification, and operator-level claims unresolved.

  • 6 LIMITATIONS: The study uses small tied-embedding Transformers, modular addition and division, only two primes, and a 2 × 2 factorial grid.Five-hundred-step sampling and 11–12 pairs limit timing and magnitude precision.
  • 6 LIMITATIONS: Activation replacement and Fourier projection create hybrid counterfactuals, donors preserve marginals rather than joint dependence, probes are associational, and tied input/output support remains inseparable.Path restoration localizes mediation at block-1 MLP, not a literal operator or unique route.
  • 6 LIMITATIONS: Untied readouts, denser checkpoints, other operations, a controlled third prime, larger grids, and natural-language models remain open; p = 53 is neither equivalence evidence nor an identified moderator.Controls remove different nuisances rather than randomizing memorization, and division and p = 43 are not unclipped-utility invariant.

7 CONCLUSION

The study concludes that grokking recodes an existing distributed circuit rather than switching between modules, while emphasizing protocol boundaries and heterogeneous validation.

  • 7 CONCLUSION: Grokking recodes an existing distributed circuit: attention leads at memorization, selected block-0 spectral modes become OOD-useful, and their effect flows chiefly through the final MLP.Prospective cohorts, division, alternative utilities, donor replacement, and exact path confirmation preserve this account.
  • 7 CONCLUSION: The reported conclusions rely on behavior-gated, version-controlled, disjoint-stage protocols, with failed seeds retained in behavior denominators and downstream tests fail-closed.Version control provides an auditable temporal boundary but is weaker than preregistration with an independent external custodian.
  • 7 CONCLUSION: Fresh addition and division validation used exactly two licensed tests after fixed seeds and separate Holm-family procedures were established.The final validation family was defined before outcomes existed, and new cohorts were kept separate rather than merged by seed identity.
  • 7 CONCLUSION: Division validation showed a variable effect whose interval crossed zero, limiting precision despite the frozen exact-test decision.The raw normal-approximation MDE was .2533, above the .1591 mean.

E FRESH PRIME-BY-WIDTH-BY-HEAD FACTORIAL RECORD

The fresh prime-by-width-by-head factorial tests whether the attention-minus-MLP structural contrast generalizes across architectural settings, while retaining explicit inferential and implementation audits.

  • E FRESH PRIME-BY-WIDTH-BY-HEAD FACTORIAL RECORD: The formal matrix trained 192 new runs across four architectures, 12 seed/split blocks, and two primes, with all eight cells unlocking and controls passing in 0/12 runs.Seven main cells passed in 12/12 runs; p=43 width 256/four heads passed in 11/12, and failed runs remained in the denominators.
  • E FRESH PRIME-BY-WIDTH-BY-HEAD FACTORIAL RECORD: The cellwise plot describes seed/split-pair means and intervals, with blue predicted head-dimension-32 diagonal cells contrasted against orange off-diagonal cells.The frozen inference is the within-seed factorial contrast, not the descriptive cell summaries.
  • E FRESH PRIME-BY-WIDTH-BY-HEAD FACTORIAL RECORD: The frozen p=43 diagonal was supported at .2115 bits/OOD example, whereas p=53 was not, so the predeclared cross-prime claim failed.The p=43 result had Holm p = .004883; p=53 had raw/Holm p = .057617.
  • E FRESH PRIME-BY-WIDTH-BY-HEAD FACTORIAL RECORD: All recovery and attribution launchers passed audit, including byte-identical replay trajectories, exact coalition games, and maximum Shapley-efficiency error of 1.819 × 10^-12.A renderer correction changed only report/figure key lookups after numerical files were written.

F FOURIER MECHANISM CONFIRMATION RECORD

The Fourier mechanism record freezes frequency selection before spectral outcomes, decomposes block-0 attention into exact Owen subplayers, and finds that selected degree-two modes explain most of the controlled gain.

  • F FOURIER MECHANISM CONFIRMATION RECORD: The frozen Fourier family selects five non-DC frequencies at the strict memory anchor, embeds their degree-two coordinates in a five-player game, and computes exact Owen values across 128 coalitions.Held-out ridge probes use training pairs only, while selective removal compares post-minus-pre OOD and training damage.
  • F FOURIER MECHANISM CONFIRMATION RECORD: Both primary Fourier hypotheses had raw exact values of .000244 and Holm-adjusted values of .000488.The full seed table, rather than checkpoint or grid-example counts, is the inferential unit.
  • F FOURIER MECHANISM CONFIRMATION RECORD: Selected key modes comprise 78.8% of the mean-baseline group contrast and 91.5% under zero replacement, with a .3808 key-minus-residual contrast positive in 12/12 pairs.These are exact additive Owen-group shares; nonlinear restricted/excluded-logit metrics are not additive clean-performance shares.
  • F FOURIER MECHANISM CONFIRMATION RECORD: Top-3 and top-8 sensitivity runs also passed in 12/12 pairs, with aggregate key shares ranging from 62.9% to 92.1%.They were secondary analyses and did not enlarge or retest the primary top-5 Holm family.
  • F FOURIER MECHANISM CONFIRMATION RECORD: Reconstruction, Owen-efficiency, archived-value, and inverse-DFT discrepancies remained below frozen tolerances across 24 manifests and 48 checkpoint payloads.Maximum Owen group-efficiency error was 8.882 × 10^-16.

G UTILITY, REPRESENTATION, AND SELECTIVE-MODE RECORD

The utility and representation record rejects the proposed MLP-memory/attention-generalization handoff and finds descriptive timing differences without support for routing onset or global rank collapse.

  • G UTILITY, REPRESENTATION, AND SELECTIVE-MODE RECORD: The attempted MLP-memory signature reverses: H-U1 is −.331 bits/example at the memory anchor, fails its zero-baseline safeguard, and remains negative at first memory and generalization.The predicted direction occurred in 0/12 pairs.
  • G UTILITY, REPRESENTATION, AND SELECTIVE-MODE RECORD: Mean half-rise locations were .570 for the block-0 attention probe, .520 for the key-degree-two spectral fraction, and .611 for block-0 attention Shapley utility.Probe-minus-utility timing was −.042, while spectrum-minus-utility timing was −.091.
  • G UTILITY, REPRESENTATION, AND SELECTIVE-MODE RECORD: These timing and boundary comparisons are descriptive and cannot rescue H-U1 or enlarge the Holm family.The spectrum-minus-utility interval was [−.170, −.012].
  • G UTILITY, REPRESENTATION, AND SELECTIVE-MODE RECORD: All representation and dense layer-game manifests completed, with maximum Fourier reconstruction error 1.192 × 10^-7 and Shapley-efficiency error 1.819 × 10^-12.The audited scientific tree contained 3,236 files.

H SPECTRAL ROBUSTNESS AND DIVISION TRANSFER RECORD

The robustness record supports spectral transfer from addition to division while preserving a distributed transition pattern and a strong downstream block-1 MLP path.

  • Division transfer: .2836 same-frequency degree-two context gain and .1545 selected spectral-power gain transfer across all 11 division pairs, whereas pure selected ratio-mode gain is .0219.The pure ratio-mode interval includes zero, while other same-frequency context, residual support, restricted multiplicative-logit gain, and spectral power show stronger transfer.
  • Division transfer: 2.0067 bits/example is the H-D3 difference from removal damage rising on OOD by 2.4491 bits and on train by .4424 bits.The division validation used 11 behavior-eligible pairs, with H-D1, H-D2, and H-D3 Holm values of .001465, .004883, and .001465.
  • Exact path mediation: 2.6440 block-1 MLP contribution exceeds the .1273 sum of the other tested paths in the disjoint cohort.The block-1-MLP-minus-others means are 2.7853 unclipped, .5479 by accuracy, and .9613 by restricted logit.
  • Dense trajectories: + .0365 attention-minus-MLP half-rise difference under mean replacement means attention reaches half-rise later, so the result does not support earlier-attention onset.Zero replacement gives +.0334, with only 8/12 positive pairs, making the ordering non-robust to baseline.
  • Control robustness: 12/12 equal-data controls were memory-only at the generalization step and had no persistent transition, supporting the paired-control contrast.The controls matched examples, architecture, initialization, and optimizer, changing only weight decay from 1 to 0.

L POST-REVIEW ROBUSTNESS AND MODERATOR RECORD

Post-review analyses preserve the central addition allocation but qualify several structural claims, while moderator checks remain descriptive and the strict anchor required correction.

  • Utility sensitivity: 78.8%, 73.0%, and 80.8% are the selected-subspace shares under clipped, unclipped, and accuracy allocations, respectively.The central addition allocation and handoff results are directionally stable, but prospective division, factorial, and restricted-logit accuracy results are not utility-invariant.
  • Scope boundary: The prospective division structural result and factorial result are not utility-invariant, despite stable central addition allocation and handoff directions.Review-v2 rescored archived exact coalition payloads without retraining models or changing frozen families.
  • Moderator checks: .0091 and .0177 within-prime Spearman correlations do not support density or single-seed explanations for factorial effects, without identifying a causal moderator.Factorial means were [.1919, .2412] for one prime and [.0143, .0254] for another after leave-one-out checks.
  • Protocol boundaries: The low-β2 test omits control subtraction and applies only among gated low-β2 runs, limiting its interpretation.Controls use the transition run’s optimizer steps and therefore need not exhibit a transition themselves.
Loading 2609.17571v1…