Page-EntroKV: Hardware-Aligned, Entropy-Weighted KV-Cache Eviction under Grouped-Query Attention
Page-EntroKV evicts KV cache entries at the same layer-group-page granularity used by GQA serving engines.
The paper argues that per-head eviction wastes memory under grouped-query attention because shared KV buffers must keep the union of different heads’ choices. Its method weights heads using sink-isolated collision entropy, then projects pooled scores onto PagedAttention frames for eviction. In pilot tests on Qwen2.5-1.5B-Instruct, it reports exact 1.000 union overhead, exact retained cardinality across page sizes, and 100% needle recall where mean pooling got 0% at a 20% budget. ArXiv · AI/CL/LG's note
The paper argues that per-head eviction wastes memory under grouped-query attention because shared KV buffers must keep the union of different heads’ choices. Its method weights heads using sink-isolated collision entropy, then projects pooled scores onto PagedAttention frames for eviction. In pilot tests on Qwen2.5-1.5B-Instruct, it reports exact 1.000 union overhead, exact retained cardinality across page sizes, and 100% needle recall where mean pooling got 0% at a 20% budget. ArXiv · AI/CL/LG's note
score 5