Which Eviction Policy Should an LLM Cache Use? A Systematic Study Across Workloads, Capacities, and Encoders
LFU was the strongest simple eviction default, and the bigger risk was whether cache hits were valid answers at all.
The study compared several LLM semantic-cache eviction policies across corpora, capacities, and encoders, and none beat LFU by more than 0.041 percentage points in any setting. FIFO and the streaming SISO adaptation fell behind LFU by more than eight points at tight capacity. The authors argue geometry-aware eviction had little room to help because newly inserted entries lacked resident neighbors within the hit radius under the tested protocol. Their audit found that many apparent semantic-cache hits were not answer-substitutable, cutting raw 51-60% hit rates to quality-adjusted rates of 1.1-2.2%. ArXiv · AI/CL/LG's note
The study compared several LLM semantic-cache eviction policies across corpora, capacities, and encoders, and none beat LFU by more than 0.041 percentage points in any setting. FIFO and the streaming SISO adaptation fell behind LFU by more than eight points at tight capacity. The authors argue geometry-aware eviction had little room to help because newly inserted entries lacked resident neighbors within the hit radius under the tested protocol. Their audit found that many apparent semantic-cache hits were not answer-substitutable, cutting raw 51-60% hit rates to quality-adjusted rates of 1.1-2.2%. ArXiv · AI/CL/LG's note
score 4