Megadose AI progress, ranked daily.

Prefix Sliding for efficient test-time scaling

· ArXiv · AI/CL/LG ·
The paper says models can drop most older reasoning tokens without losing the thread.

Prefix Sliding keeps the initial instructions and tools plus only the most recent reasoning window, cutting memory use during long test-time reasoning. The authors report existing models can run about 3x faster without training while maintaining performance. With reinforcement learning under the same constraint, they say models can scale to reasoning traces beyond 100,000 tokens and improve results. The paper also says Prefix Sliding beats both summarizing old tokens and a plain sliding-window setup. ArXiv · AI/CL/LG's note

score 6

Categories: Research