Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV
Retained KV rows can preserve a missing observation well enough to steer later answers.
The paper tests sparse KV-cache memory by removing an earlier observation from served context while keeping later cached events. In sensitive cases, answers still largely follow the omitted value, which the authors call semantic materialization. A deliberately phrased, answer-free event lifted donor-aligned recovery on Qwen3-8B from 6% to 51%, while natural long-dialog mentions showed no detected advantage. The effect is bounded: compact state survives better than larger payloads, and wording can determine whether the cache writes useful memory at all. HF Daily Papers' note
The paper tests sparse KV-cache memory by removing an earlier observation from served context while keeping later cached events. In sensitive cases, answers still largely follow the omitted value, which the authors call semantic materialization. A deliberately phrased, answer-free event lifted donor-aligned recovery on Qwen3-8B from 6% to 51%, while natural long-dialog mentions showed no detected advantage. The effect is bounded: compact state survives better than larger payloads, and wording can determine whether the cache writes useful memory at all. HF Daily Papers' note
score 5