KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems
The paper’s claim is that multi-agent systems can share KV caches without paying for separate prefills per agent.
KVCMAS corrects cache differences caused by agent-specific prefixes using compact low-rank states. It chains those corrections through the agent workflow, avoiding an extra reference prefill while keeping the first agent’s cache exact. In the reported tests, it matched or improved prior accuracy while cutting time-to-first-token under concurrent serving. The authors report a 2.0x TTFT speedup over no cache sharing and up to 3.7x lower peak GPU memory than a prior correction method. HF Daily Papers' note
KVCMAS corrects cache differences caused by agent-specific prefixes using compact low-rank states. It chains those corrections through the agent workflow, avoiding an extra reference prefill while keeping the first agent’s cache exact. In the reported tests, it matched or improved prior accuracy while cutting time-to-first-token under concurrent serving. The authors report a 2.0x TTFT speedup over no cache sharing and up to 3.7x lower peak GPU memory than a prior correction method. HF Daily Papers' note
score 4