Megadose AI progress, ranked and analyzed.

LoGo: Token-Level Dynamic Local-Global Attention

· ArXiv · AI/CL/LG ·
LoGo gives every token local attention, then selectively grants full-context attention to tokens a learned gate flags as needing it.

The paper frames attention span as the budget knob for long-context models. Its controller targets a fixed global-attention ratio without auxiliary losses, while progressive masking is used to stabilize training before sparse routing begins. The authors say query-sparse Triton kernels turn the reduced global computation into practical speedups. In experiments, LoGo is reported to preserve full-attention scaling behavior and beat both full-attention Transformers and static local-global hybrids in controlled comparisons, with stronger long-range retrieval results. ArXiv · AI/CL/LG's note

score 5

Categories: Research