Periscope: Extending Frozen Language Models Beyond Their Context Window
Periscope turns long-context reading into a grid of short probes, letting frozen models score evidence beyond their normal window.
The method splits a text into chunks, places them on a square grid, and asks the same finite-choice question over local and strided spans. It uses the resulting log-odds to build an evidence map, then reads only the highest-ranked chunks. The paper says this matched or beat long-window reads on LongBench v2 and InfiniteBench while using much less KV cache. A 27B model is reported reading 4.5M-token contexts on one 80GB GPU.
HF Daily Papers' note
The method splits a text into chunks, places them on a square grid, and asks the same finite-choice question over local and strided spans. It uses the resulting log-odds to build an evidence map, then reads only the highest-ranked chunks. The paper says this matched or beat long-window reads on LongBench v2 and InfiniteBench while using much less KV cache. A 27B model is reported reading 4.5M-token contexts on one 80GB GPU.
HF Daily Papers' note
score 4