Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position
Hybrid attention designs show a tradeoff between context extension and length extrapolation.
The paper compares hybrids using full attention with sliding-window attention or gated linear attention variants. It reports a “Seesaw Effect”: linear-attention hybrids benefit more from long-context continual pretraining, while sliding-window hybrids do better under length extrapolation. The authors attribute the split to different positional inductive biases, and propose Sliding-Window Linear Attention as a way to get 16x training-free length extrapolation while preserving 100% NIAH-SK1 accuracy at 64k context. HF Daily Papers' note
The paper compares hybrids using full attention with sliding-window attention or gated linear attention variants. It reports a “Seesaw Effect”: linear-attention hybrids benefit more from long-context continual pretraining, while sliding-window hybrids do better under length extrapolation. The authors attribute the split to different positional inductive biases, and propose Sliding-Window Linear Attention as a way to get 16x training-free length extrapolation while preserving 100% NIAH-SK1 accuracy at 64k context. HF Daily Papers' note
score 5