CoWindow Attention: Full Causal Coverage Is a Collective Property
CoWA claims near-FullAttn results by splitting long-range causal history across heads instead of repeating it in every head.
The paper says all heads keep local and prefix-sink windows, while complementary long-range windows divide the rest of the context. In an 8K ablation, CoWA with full collective coverage reached 89.73% accuracy, close to FullAttn’s 89.97%, while duplicated long-range windows did worse. At 128K tokens, the authors report 7.4x/8.6x lower training forward/backward latency and 3.0x lower decoding latency versus FullAttn. Scaling runs from 0.6B to 14B, plus separate 14B and 32B continued-training models, are reported as broadly comparable on perplexity and downstream scores. HF Daily Papers' note
The paper says all heads keep local and prefix-sink windows, while complementary long-range windows divide the rest of the context. In an 8K ablation, CoWA with full collective coverage reached 89.73% accuracy, close to FullAttn’s 89.97%, while duplicated long-range windows did worse. At 128K tokens, the authors report 7.4x/8.6x lower training forward/backward latency and 3.0x lower decoding latency versus FullAttn. Scaling runs from 0.6B to 14B, plus separate 14B and 32B continued-training models, are reported as broadly comparable on perplexity and downstream scores. HF Daily Papers' note
score 5