Megadose AI progress, ranked daily.

Prefix Sliding for efficient test-time scaling

· HF Daily Papers ·
The method keeps the prompt prefix and recent reasoning, then drops the middle to cut memory costs.

Prefix Sliding targets the cost of long test-time reasoning, where full attention keeps every intermediate token in memory. The paper says many middle reasoning tokens become less important as generation continues, so the model retains key instructions/tools plus the latest working context. Without training, the authors report up to 3x faster existing models while maintaining performance. With reinforcement learning, they say it can scale reasoning traces past 100,000 tokens and outperform summarization or a plain sliding window. HF Daily Papers' note

score 6

Categories: Research