MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution
MS-GLA splits gated linear attention heads across temporal scales to reduce the memory bottleneck in fixed-size recurrent attention.
The paper argues that standard GLA forces every head to handle token-level syntax and long-range semantics at one resolution. MS-GLA assigns some heads to finer spans and others to coarser pooled spans, then uses a learnable input-dependent fusion layer to recombine them at each step. The authors report gains over GLA at matched parameter counts, including up to 18.9% on recall-heavy tasks and 9.5% lower average perplexity on language modeling benchmarks. Source: ArXiv · AI/CL/LG's note.
The paper argues that standard GLA forces every head to handle token-level syntax and long-range semantics at one resolution. MS-GLA assigns some heads to finer spans and others to coarser pooled spans, then uses a learnable input-dependent fusion layer to recombine them at each step. The authors report gains over GLA at matched parameter counts, including up to 18.9% on recall-heavy tasks and 9.5% lower average perplexity on language modeling benchmarks. Source: ArXiv · AI/CL/LG's note.
score 5