Megadose Built for builders and researchers.

Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling

· HF Daily Papers ·
The paper proposes a linear-attention variant that stores history in a 3D tensor state instead of a matrix.

Triadic linear attention writes a triadic outer product of two keys and a value into that state, then reads it with two queries. The authors say the added second key can multiply state size by `E` while requiring only two extra projections. They report compatibility with data-dependent forgetting, the delta rule, and chunkwise-parallel training. Applied to Gated DeltaNet and scalar-gated linear attention, it improves long-context language modeling and recall against other state-enlarging alternatives. HF Daily Papers' note

score 5

Categories: Research