Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
The paper says hybrid linear-attention LLMs develop massive-activation spikes right before full-attention layers, with plateaus that can carry through linear-attention blocks.
The authors report this pattern across five linear-attention architectures, six hybrid setups, five data domains, and open-source hybrid models from 1.2B to 397B parameters. In controlled pretraining up to 1.3B parameters, the spike and plateau shapes appear early. Full-attention output gating reduces their magnitude but does not remove the layerwise structure. The proposed mechanism is a timing difference in cancellation: localized for pre-attention spikes, delayed for inter-spike plateaus. ArXiv · AI/CL/LG's note.
The authors report this pattern across five linear-attention architectures, six hybrid setups, five data domains, and open-source hybrid models from 1.2B to 397B parameters. In controlled pretraining up to 1.3B parameters, the spike and plateau shapes appear early. Full-attention output gating reduces their magnitude but does not remove the layerwise structure. The proposed mechanism is a timing difference in cancellation: localized for pre-attention spikes, delayed for inter-spike plateaus. ArXiv · AI/CL/LG's note.
score 5