PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction
The paper claims multi-agent LoRA systems can share KV cache without training while keeping each agent’s adapter behavior intact.
PReCache splits the shared context cache into a base cache and compact agent-specific low-rank caches, avoiding repeated prefill over long shared trajectories. Its PreLRShared design precomputes each agent’s low-rank cache the first time the context is processed. ReBaseShared reconstructs the shared base cache from adapter-free hidden states to reduce errors from reusing another agent’s adapted cache. The authors report up to 3.1x faster TTFT, 2.3x better per-request throughput, and an average accuracy drop of 1.1 points for ReBaseShared versus no cache sharing. HF Daily Papers' note
PReCache splits the shared context cache into a base cache and compact agent-specific low-rank caches, avoiding repeated prefill over long shared trajectories. Its PreLRShared design precomputes each agent’s low-rank cache the first time the context is processed. ReBaseShared reconstructs the shared base cache from adapter-free hidden states to reduce errors from reusing another agent’s adapted cache. The authors report up to 3.1x faster TTFT, 2.3x better per-request throughput, and an average accuracy drop of 1.1 points for ReBaseShared versus no cache sharing. HF Daily Papers' note
score 4