Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating
The paper says its memory-eviction method is a framework result, not a new benchmark winner.
It treats test-time memory eviction as a fixed-lag estimation problem: wait briefly, see what the model actually used, then decide what to keep. The authors introduce RMM, a training-free policy that generalizes H2O and works well in controlled settings where reuse is delayed and endogenous. On NVIDIA’s KVPress benchmarks, though, the gain mostly vanishes; RMM matches H2O on single-turn QA and loses to H2O and SnapKV in streaming multi-turn tests. ArXiv · AI/CL/LG's note
It treats test-time memory eviction as a fixed-lag estimation problem: wait briefly, see what the model actually used, then decide what to keep. The authors introduce RMM, a training-free policy that generalizes H2O and works well in controlled settings where reuse is delayed and endogenous. On NVIDIA’s KVPress benchmarks, though, the gain mostly vanishes; RMM matches H2O on single-turn QA and loses to H2O and SnapKV in streaming multi-turn tests. ArXiv · AI/CL/LG's note
score 5