Megadose AI progress, ranked and analyzed.

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

· HF Daily Papers ·
Random eviction worked about as well as scored KV-cache pruning once the prompt was protected.

The paper says Random Attention keeps the prompt, then evicts cached tokens uniformly at random within each attention head. Across four models and six reasoning tasks, it matched the strongest prior eviction method while delivering 32-43% higher throughput in vLLM deployment. The authors argue the prompt is the cache’s fragile part, while reasoning traces are redundant enough in text and across heads to survive random pruning. Code is described as publicly available. HF Daily Papers' note

score 6

Categories: Research