Megadose Built for builders and researchers.

The Extender: A Log-Structured Transformer

· HF Daily Papers ·
A second, append-only channel carries the attention keys and values, cutting persistent attention memory sharply.

Jakob Eriksson’s paper adds a small concatenation stream, `x`, alongside the usual Transformer residual stream. Attention `kv` projections read from that appended stream, while the FFN and queries still see the residual state. With 32-dimensional extensions, the Extender matches Transformer accuracy on short-context CORE tasks at 199M-924M parameters and beats it on long-context RULER workloads at 924M. The 1664-wide, 924M model is reported to use 104x less persistent attention memory than MHA. HF Daily Papers' note

score 5

Categories: Research