Megadose AI progress, ranked and analyzed.

Which the Eye Fears: Writing with Read-Blindness Explains Massive Activations in Transformers

· ArXiv · AI/CL/LG ·
The paper argues massive activations persist because transformer blocks are blind to those coordinates when reading, while still writing into them.

The authors say this read-write asymmetry appears in both attention and feed-forward layers, letting extreme residual-stream features accumulate without corrective feedback. Checkpoint analysis finds read-blindness emerges before feed-forward amplification, putting it upstream of an earlier proposed mechanism. Gradient analysis is used to argue the model maintains the asymmetry rather than merely failing to remove it. Attempts to remove read-blocking in one place trigger compensating shifts elsewhere, and the massive activations remain. ArXiv · AI/CL/LG's note

score 5

Categories: Research