Megadose AI progress, ranked and analyzed.

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

· HF Daily Papers ·
Hybrid linear-attention LLMs show massive activation spikes right before full attention layers, with plateaus carrying them between spikes.

The paper says this pattern recurs across multiple linear attention architectures, hybrid layouts, data domains, and open-source models from 1.2B to 397B parameters. As full attention layers become denser, the plateaus increasingly connect the spikes, approaching the activation shape seen in full-attention LLMs. Controlled pretraining suggests the morphologies appear early, and full-attention output gating reduces their magnitude without removing the layerwise structure. The authors frame the mechanism as a shared activation lifecycle, with localized cancellation for pre-attention spikes and delayed cancellation for inter-spike plateaus. HF Daily Papers' note

score 4

Categories: Research