Megadose Built for builders and researchers.

Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position

· HF Daily Papers ·
Hybrid attention designs show a tradeoff between context extension and length extrapolation.

The paper compares hybrids using full attention with sliding-window attention or gated linear attention variants. It reports a “Seesaw Effect”: linear-attention hybrids benefit more from long-context continual pretraining, while sliding-window hybrids do better under length extrapolation. The authors attribute the split to different positional inductive biases, and propose Sliding-Window Linear Attention as a way to get 16x training-free length extrapolation while preserving 100% NIAH-SK1 accuracy at 64k context. HF Daily Papers' note

score 5

Categories: Research