Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
Chunked KV-cache compression can make retrieval accuracy swing sharply depending on where a token falls inside the compression window.
The paper calls this “phase sensitivity”: information may be easy to retrieve at one window-relative position and much harder at another. In large open-weight models using this kind of compression, the authors report long-context retrieval gaps of up to 40 percentage points across phases. They also pretrain smaller transformers across several KV-compression designs and reproduce the effect. Their analysis points to phase-specialized attention components, meaning average benchmark scores can hide regular positional failures. HF Daily Papers' note
The paper calls this “phase sensitivity”: information may be easy to retrieve at one window-relative position and much harder at another. In large open-weight models using this kind of compression, the authors report long-context retrieval gaps of up to 40 percentage points across phases. They also pretrain smaller transformers across several KV-compression designs and reproduce the effect. Their analysis points to phase-specialized attention components, meaning average benchmark scores can hide regular positional failures. HF Daily Papers' note
score 5