KV$^2$: A Self-Refining KV Cache
KV² compresses reusable long-context KV caches by rescoring only selected tokens, not the full prompt.
The paper targets cases where one prefilled context is reused for many later queries, making KV-cache memory a main cost. KV² uses a lightweight proxy to pick informative in-context tokens, then reprocesses only that subset to decide what to evict. In the authors’ tests, its advantage grows at tighter cache budgets, including a more than 40-point gain over the next-best baseline on RULER 16K at a 2% budget. It also reports the best LongBench average across 2%-10% budgets while using less compression-stage runtime and peak memory than full-context reconstruction. ArXiv · AI/CL/LG's note
The paper targets cases where one prefilled context is reused for many later queries, making KV-cache memory a main cost. KV² uses a lightweight proxy to pick informative in-context tokens, then reprocesses only that subset to decide what to evict. In the authors’ tests, its advantage grows at tighter cache budgets, including a more than 40-point gain over the next-best baseline on RULER 16K at a 2% budget. It also reports the best LongBench average across 2%-10% budgets while using less compression-stage runtime and peak memory than full-context reconstruction. ArXiv · AI/CL/LG's note
score 5