Megadose Built for builders and researchers.

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

· HF Daily Papers ·
HLA lets each query choose which compressed history chunks to use, instead of relying on fixed chunk mixing.

The paper adds query-dependent chunk-level routing to Gated DeltaNet, representing completed chunks as compact affine state transitions. Those gates can preserve or suppress a chunk’s memory and its effect on earlier recurrent state. In tests on Qwen3.5 models from 0.8B to 9B, HLA beats native GDN and fixed chunk mixing, including gains up to 5.57 points on LongBench-V2 and 3.97 on RULER. A from-scratch 1.3B model trained on 100B tokens also improves on RULER beyond its 4K training context, with larger gains at 32K. HF Daily Papers' note

score 5

Categories: Research