Source-linked AI summary

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

Zunhai Su, Bohan Sun, Xialie Zhuang, Shuibai Zhang, He Xiao, Jing Xiong, Hengyuan Zhang, Zhongzhu Zhou, Tiantian Zhang, Ngai Wong, Chuan-Wei Kuo

arXiv:2608.12149v1cs.CL

TL;DR

The paper addresses limited understanding of massive activations in layer-interleaved hybrid linear attention LLMs. It systematically characterizes their layerwise organization, training emergence, gating response, and cancellation dynamics, finding recurring pre-attention spikes and inter-spike plateaus that converge toward full-attention morphology. The results support a shared lifecycle in which cancellation timing distinguishes localized spikes from persistent plateaus.

  • Problem

    How hybrid linear attention reshapes massive activations remains poorly understood, despite their value as a probe of layerwise computation.

  • Method

    The study combines inference-time analyses across architectures, hybridization configurations, scales, and domains with controlled GDN pretraining and systematic-outlier tracking.

  • Results

    Massive activations consistently form pre-attention spikes and, with denser full attention, inter-spike plateaus; both emerge early, while full attention gating attenuates magnitudes without eliminating organization.

  • Takeaways & Limitations

    PAS and ISP are recurring architecture-aligned morphologies explained by a shared cancellation-timing lifecycle, with delayed cancellation sustaining inter-spike plateaus.

  • Takeaways & Limitations

    The qualitative cross-architecture comparison fixes model scale at 1.3B parameters and the hybridization ratio at 12:1 while varying only the linear attention mechanism.

Abstract

from arXiv · show

We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter-spike plateaus (ISP). As full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology of full attention LLMs. We establish the recurrence of this organization across five linear attention architectures, six hybridization configurations, five data domains, and representative open-source hybrid models spanning 1.2B to 397B total parameters. Controlled pretraining of GDN-based hybrids at scales up to 1.3B shows that both morphologies emerge early and respond asymmetrically to output gating: full attention output gating strongly attenuates their absolute magnitudes without eliminating their layerwise organization, whereas removing GDN gates yields comparatively modest amplification. Mechanistically, our systematic-outlier analysis supports a shared lifecycle account governed by the timing of MA cancellation. PAS follows a localized write-sink-cancel process, while the extended persistence of ISP is consistent with delayed cancellation. At the full attention limit, this account recovers the stable MA morphology characteristic of full attention LLMs. Our code is available at https://github.com/StartluxLabs/Massive-Activations-HLA.

1 INTRODUCTION

The paper studies how hybrid linear attention reorganizes massive activations, identifying pre-attention spikes and inter-spike plateaus as architecture-aligned morphologies. These patterns recur across architectures, configurations, scales, and domains, while cancellation timing provides a shared mechanistic account.

  • Research gap and contribution: Hybrid linear attention reshapes massive activations into pre-attention spikes before full attention layers and inter-spike plateaus across intervening linear layers.At the full attention limit, the spikes and plateaus merge into the stable morphology seen in full attention LLMs.
  • Research gap and contribution: The study systematically characterizes massive activations across five linear attention architectures, six hybridization configurations, and five data domains.The analysis targets the poorly understood behavior of massive activations beyond conventional full attention models.
  • Training dynamics: Controlled pretraining shows that PAS and ISP emerge early and consolidate during optimization.The experiments use GDN-based hybrids at scales up to 1.3B parameters.
  • Training dynamics: Full attention output gating strongly attenuates PAS and ISP magnitudes without removing their layerwise organization, whereas removing GDN gates causes comparatively modest amplification.The two gating interventions therefore affect absolute magnitude asymmetrically.
  • Mechanistic account: A shared systematic-outlier lifecycle links PAS to localized write–sink–cancel dynamics and ISP to delayed cancellation.The same account recovers the stable full-attention morphology at the full attention limit.

2 PRELIMINARIES ON HYBRID LINEAR ATTENTION MODELS

Hybrid linear attention interleaves full and linear attention layers to combine attention’s modeling capacity with recurrent efficiency. The section defines the hybridization ratio and contrasts full attention’s quadratic cost with linear attention’s fixed-size recurrent state.

  • Hybrid architecture: An HLA model uses full attention at selected layers and linear attention elsewhere within pre-normalized residual blocks.The residual-stream representation after block ℓ is denoted X^(ℓ).
  • Hybrid architecture: The hybridization ratio ρ = L/L_FA specifies one full attention layer per ρ sequence-mixing layers; larger ρ means sparser full attention, while ρ = 1 recovers full attention.L_FA is the number of full attention layers.
  • Attention mechanisms: Full attention enables direct, content-dependent interactions with all preceding tokens but incurs quadratic complexity in sequence length.This cost motivates combining full attention with more efficient sequence-mixing mechanisms.
  • Attention mechanisms: Linear attention replaces pairwise token interactions with a fixed-size recurrent state, enabling linear-time processing without a growing KV cache.Its recurrent state size is independent of sequence length, although fixed state can restrict modeling capacity for recall-intensive tasks.
  • Motivation: Interleaving full and linear attention combines complementary sequence-mixing mechanisms whose interaction may reorganize internal computation across layers.This unresolved interaction motivates studying layerwise massive-activation organization.

3 MASSIVE ACTIVATION DYNAMICS IN HYBRID LINEAR ATTENTION MODELS

The study tracks massive activations across hybrid linear attention models and finds architecture-aligned pre-attention spikes and inter-spike plateaus that recur across architectures, scales, hybridization ratios, and domains.

  • Experimental setup: The analysis tracks massive activations using attention-sink-guided token identification across five linear attention architectures, multiple scales, ratios, and five input domains.The evaluation includes 340M and 1.3B models and compares hybrid configurations with pure linear and full attention limits.
  • Attention-sink-guided tracking: Attention sinks show less stable magnitude rankings in HLA models, so the study complements magnitude ranking with attention-derived consensus sink anchors.The maximally activated token also switches more frequently between adjacent layers than in the full attention baseline.
  • Pre-attention spikes: Across all five linear attention backbones, sink-associated activations form pronounced local maxima immediately before full attention layers, termed pre-attention spikes (PAS).Macro-average PAS localization remains consistently high across all ten architecture–scale pairs and exceeds non-sink controls.
  • Inter-spike plateaus: PAS persist across hybrid configurations, while denser full attention progressively elevates intervening activations into inter-spike plateaus (ISP).At the full attention limit, the distinction between spikes and plateaus disappears, recovering a model-wide plateau morphology.
  • Cross-model recurrence: PAS and ISP recur across large-scale hybrid models, sequence mixers, model scales, and domains, with layerwise positions more stable than activation magnitudes.Matched checkpoints retain aligned PAS locations and ISP boundaries, while magnitudes vary across scales, domains, and post-training stages.
  • Controlled pretraining: Full attention output gating produces larger changes than removing native GDN gates, attenuating but not eliminating the architecture-aligned organization.The disproportionate response indicates a central organizing role for full attention, while GDN gating mainly modulates propagation.

4 A SYSTEMATIC-OUTLIER ACCOUNT OF PAS AND ISP

The analysis explains PAS and ISP through a shared systematic-outlier lifecycle governed by cancellation timing. PAS is a localized write–sink–cancel event, whereas ISP reflects delayed cancellation across intervening linear-attention layers and connects toward the full-attention morphology.

  • A Systematic-Outlier Account of PAS: PAS is a coordinated cross-layer event in which pronounced massive activations arise immediately before full attention layers.The recurrence is reported across linear-attention architectures and hybridization configurations.
  • Pre-Attention Spikes: Localized Write–Sink–Cancel: Fixed-coordinate tracing identifies three PAS stages: an extreme residual-stream write, attention-sink coupling during full attention, and later opposite-signed cancellation.The cancellation substantially reduces the outlier and causes the spike to dissipate.
  • Inter-Spike Plateaus: ISP extends the same lifecycle across greater depth, with an outlier written before full attention, retained through intervening linear layers, and later canceled by an opposite-signed update.This delayed-cancellation pattern is illustrated as a plateau that eventually dissipates.
  • Connection to the Full Attention Limit: As full attention becomes denser, progressively deferred cancellation connects successive PAS through ISP and recovers the stable massive-activation morphology of full-attention models.The interpretation links hybridization-dependent morphology to cancellation timing.

5 CONCLUSION

The paper identifies PAS and ISP as recurring, architecture-aligned massive-activation morphologies in layer-interleaved HLA LLMs. It interprets them through cancellation timing, while controlled pretraining shows early emergence and attenuation by full-attention output gating.

  • 5 CONCLUSION: The study identifies PAS and ISP across evaluated architectures, hybridization configurations, model scales, and input domains.Both morphologies emerge early during controlled pretraining.
  • 5 CONCLUSION: Full-attention output gating attenuates massive-activation magnitudes but does not eliminate their layerwise organization.The conclusion contrasts magnitude attenuation with persistence of organization.
  • 5 CONCLUSION: Prompt cancellation localizes massive activations as PAS, whereas delayed cancellation sustains them across layers as ISP and recovers the full-attention morphology at the limit.The paper presents this as a shared systematic-outlier lifecycle distinguished by cancellation timing.

6 ETHICS STATEMENT

The supplied passages describe linear-attention architectures, hybridization configurations, pretrained checkpoint coverage, and experimental analyses of massive-activation dynamics. They also state that the research followed ethical standards and used publicly available or appropriately permitted data.

  • Ethics Statement: The research states that experiments used publicly available or appropriately permitted data and no sensitive or personally identifiable information.It also states that reported results and scientific claims were independently verified and approved by the authors.
  • Background: Linear attention uses recurrent state updates and fixed-size states for linear-time sequence processing.Representative families include RetNet, HGRN, GLA, DeltaNet, and GDN.
  • Model Suite and Pretraining: The inference-time analysis covers five linear-attention architectures, two parameter scales, multiple hybrid configurations, and full-attention references.The checkpoints are publicly available, and the shared pretraining setup reduces variation unrelated to architecture or attention density.
  • Experimental Design: The model suite varies full-attention density while holding each linear-attention backbone fixed, enabling controlled comparisons of massive-activation dynamics.Configurations span hybrid ratios from 24:1 to 3:1, plus full-attention and pure-linear-attention limits.
  • Evaluation: Across representative inputs and domains, sink-associated activations show qualitatively similar layerwise organization, with maxima immediately before full-attention layers.The analysis includes WikiText-103, Scientific Papers, GSM8K, CodeSearchNet, FLORES-200, and a running example.

B.5 STATISTICAL CONTROLS AND UNCERTAINTY

Statistical controls show that PAS is concentrated at sink-associated positions and that denser full attention increases inter-spike persistence. Absolute plateau activation generally rises as well, while the normalized and unnormalized measures capture different properties.

  • Uncertainty: 10,000 stratified bootstrap resamples quantify alignment and ISR uncertainty across the same 500 inputs and five domains.Inputs are sampled with replacement within each domain before recomputing macro-averages.
  • Sink–spike controls: 6.5–58.6 percentage-point sink-minus-non-sink alignment gaps and 0.97–4.42 log2-unit peak-excess gaps exclude zero across all ten architecture–scale pairs.Every paired 95% bootstrap confidence interval excludes zero.
  • ISR uncertainty: 5.5–44.5 and 17.0–52.3 percentage-point increases across adjacent hybridization ratios confirm that ISR rises with denser full attention.All 20 architecture–scale comparisons increase, with paired 95% confidence intervals excluding zero.
  • Absolute plateau activation: 19 of 20 adjacent-ratio comparisons show increased mean absolute inter-spike activation as full attention becomes denser.The normalized ISR measures relative persistence, whereas AISP measures absolute plateau amplitude.
  • Cross-model scope: The cross-model evaluation spans 12 checkpoints, four hybrid families, and 1.2B–397B total parameters, covering periodic and nonuniform hybridization.The models include linear-attention and state-space mixers, with base and instruction-tuned checkpoints.

C.2 MA DYNAMICS ACROSS MODELS AND INPUT DOMAINS

Across hybrid model families, scales, post-training stages, mixers, and representative domains, PAS and ISP remain aligned with full attention placement. Layerwise morphology is more stable than activation magnitude.

  • Post-training stages: PAS locations and ISP spans remain closely aligned between matched base and instruction-tuned Kimi Linear and Qwen3.5 checkpoints.Activation magnitudes differ, but the correspondence with full attention placement remains largely intact.
  • Model sizes: Across Qwen3.5, Nemotron-H, and Zamba2 sizes, PAS and ISP stay aligned with full attention placement despite varying magnitude and prominence.Magnitude has no consistent monotonic relationship with parameter count.
  • Sequence mixers: PAS and ISP recur in both linear-attention and state-space hybrids, with spikes before full attention and plateaus extending through intervening sequence-mixing blocks.This consistency supports an association with layer-interleaved hybridization rather than a particular mixer.
  • Input domains: Across six representative input types, domain variation primarily changes activation magnitude while PAS locations and ISP spans remain closely aligned.The inputs include a running example and five domains differing in language, structure, and semantic content.
  • Overall pattern: Full attention placement is the most stable organizer of PAS locations and ISP extent across the evaluated families, sizes, stages, mixers, and domains.The evidence extends beyond the controlled M-A-P suite.

D.1 EXPERIMENTAL SETUP AND ANALYSIS OVERVIEW

The controlled study uses fixed-depth GDN hybrids to isolate full attention placement and density, then traces activation formation, gating, and benchmark behavior. Training and evaluation cover language modeling, commonsense reasoning, and retrieval.

  • Model architecture: 24-layer GDN hybrids at 340M and 1.3B replace selected mixers with full attention while holding other architectural components fixed.This isolates effects of full attention placement and density from depth, width, and feed-forward architecture.
  • Training recipe: Models are trained from scratch on FineWeb-Edu with AdamW, 4,096-token sequences, gradient clipping, warmup, and cosine learning-rate decay.The 340M models use a 10B-token default training budget unless otherwise specified.
  • Analysis overview: The analysis first studies PAS across placements and scales, then ISP formation, and finally full-attention and GDN output-gating interventions.The interventions are designed to isolate output-gating effects.
  • Evaluation: Evaluation covers WikiText-103 perplexity, zero-shot commonsense accuracy, real-world retrieval, and synthetic single-needle retrieval tasks.Retrieval benchmarks include SWDE, FDA, SQuAD, NQ, and three RULER NIAH variants.
  • Representative configurations: 1.3B PAS and ISP configurations are trained on 50B tokens, comparing one layer-12 full-attention layer with eight full-attention layers under a 3:1 ratio.Figure 8 reports training loss for these representative configurations.
  • PAS measurement: PAS formation is measured by tracing maximum absolute first-token hidden-state activation across depth at successive training checkpoints.The setup compares full-attention insertion at layers 4, 12, and 20 in 340M models.

D.3 TRAINING-TIME EMERGENCE OF PAS

Controlled training shows that PAS emerges early, strengthens with deeper full-attention placement, and recurs at 1.3B scale. Placement also changes retrieval performance despite near-perfect sink–spike alignment.

  • Placement depth: PAS localizes immediately before the inserted full-attention layer, with layer 20 producing the strongest spike and layer 4 only a weak local maximum.Layer 12 produces a pronounced intermediate spike, establishing a strong dependence on placement depth.
  • Output gating: Full-attention output gating attenuates the PAS spike without eliminating it, while removing native GDN output gates moderately increases its magnitude.Figure 11 traces first-token maximum absolute activation across depth at successive 1B-token checkpoints.
  • Training-time emergence: PAS is visible after 1B tokens at 340M and reaches pronounced, persistent localization by 10B tokens at 1.3B.The 1.3B trajectory shows a weak precursor at 5B tokens and remains localized through 50B tokens.
  • Cross-scale recurrence: 100% final sink–spike alignment occurs at both 340M and 1.3B despite differences in model scale, training budget, and activation profile.This indicates recurrence of the architecture-aligned position across the tested scales.
  • Downstream performance: Middle and late full-attention placements outperform early placement on most real-world and synthetic retrieval benchmarks despite comparable alignment and language-modeling performance.All three configurations achieve near-perfect alignment.

D.4 TRAINING-TIME EMERGENCE OF ISP

Controlled pretraining shows that ISP emerges early, consolidates over training, and recurs across the two evaluated model scales. Output gating changes absolute magnitudes asymmetrically while preserving the layerwise organization of PAS and ISP.

  • Training-time emergence: ISP is already visible after 1B training tokens and progressively consolidates into a sustained plateau during pretraining.The elevated region between adjacent PAS becomes more pronounced and stable across checkpoints.
  • Cross-scale recurrence: 94.47% ISR at 1.3B versus 90.56% at 340M supports strong inter-spike retention at both evaluated scales.The comparison supports recurrence across scales without implying a controlled scaling effect.
  • Output-gating effects: Adding output gates to full attention substantially attenuates PAS and ISP magnitudes without eliminating their layerwise organization.Weak spikes and a low-magnitude plateau remain visible and persist during pretraining.
  • Output-gating effects: Removing GDN output gates increases PAS and ISP magnitudes, with a more visible effect on the plateau, but less than full-attention gating.The interventions isolate a comparatively modest GDN-gating effect.
  • Output-gating effects: Gating the fewer full attention layers produces larger trajectory changes than removing gates from all GDN layers.The results indicate that gate placement matters more than gate count for absolute MA dynamics.

E.1 FIXED-COORDINATE SYSTEMATIC-OUTLIER ANALYSIS OF PAS

Fixed-coordinate analysis identifies PAS as a localized lifecycle: an outlier is written before full attention, coupled to attention-sink behavior, and then substantially canceled.

  • Fixed-coordinate analysis: A fixed token–feature coordinate traces PAS through outlier formation, attention-sink coupling, and cancellation.The analysis follows the coordinate across the three successive stages using module-level updates.
  • Outlier Formation: The pre-attention layer writes a concentrated extreme update into the residual stream at one token–feature coordinate.This localization distinguishes PAS from a diffuse increase across tokens or features.
  • Attention-Sink Coupling: During full attention at layer 12, the corresponding sink token receives a disproportionate share of attention from subsequent query positions.The fixed-coordinate outlier therefore corresponds to sink behavior in the full attention computation.
  • Outlier Cancellation: At layer 12, a large opposite-signed update substantially cancels the incoming outlier and causes the PAS to dissipate.The cancellation sharply reduces the outlier’s residual-stream magnitude at the tracked coordinate.
  • PAS lifecycle: These stages form a coherent PAS lifecycle rather than unrelated extreme values across adjacent layers.The fixed-coordinate evidence links formation, sink coupling, and cancellation as successive processes.

E.2 CROSS-MODEL SYSTEMATIC-OUTLIER ANALYSES

Cross-model systematic-outlier analyses show that PAS and ISP remain aligned with hybrid layer placement across controlled architectures, domains, hybridization ratios, and large pretrained models. Denser full attention connects spikes through longer elevated-activation regions, approaching the sustained morphology of full attention.

  • Controlled M-A-P Models: Controlled analyses cover five linear-attention backbones and hybridization ratios of 24:1, 12:1, 6:1, and 3:1.The M-A-P suite includes GDN, DeltaNet, GLA, HGRN, and RetNet at 1.3B scale.
  • Controlled M-A-P Models: Dominant outliers arise immediately before full attention and are followed by opposite-signed updates across backbones and hybridization ratios.The recurring decomposition is consistent with delayed cancellation between successive spikes.
  • Hybridization ratios: As full attention becomes denser, elevated activations span progressively larger inter-spike intervals, transitioning from isolated PAS to sustained ISP.At the full attention limit, the localized spikes converge toward the sustained full-attention morphology.
  • Large-Scale Open-Source Hybrid Models: Across large-scale Qwen3.5 and Nemotron-H checkpoints, PAS and ISP remain aligned with hybrid layer arrangements while absolute magnitudes vary.The recurrence extends from Gated DeltaNet hybrids to Mamba-2 state-space hybrids and across model scales.
  • Large-Scale Open-Source Hybrid Models: Native full-attention output gating attenuates absolute MA magnitudes without eliminating their architecture-aligned organization.This pattern is reported in the Qwen3.5 results and contributes to the broader cancellation-timing interpretation.
  • Input domains: Across five 1.3B HLA architectures and five representative domain-specific inputs, PAS consistently emerge immediately before full attention layers.The cross-domain analysis traces the first token across depth under a fixed 12:1 hybridization ratio.
Loading 2608.12149v1…