Megadose AI progress, ranked and analyzed.

PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction

· HF Daily Papers ·
The paper claims multi-agent LoRA systems can share KV cache without training while keeping each agent’s adapter behavior intact.

PReCache splits the shared context cache into a base cache and compact agent-specific low-rank caches, avoiding repeated prefill over long shared trajectories. Its PreLRShared design precomputes each agent’s low-rank cache the first time the context is processed. ReBaseShared reconstructs the shared base cache from adapter-free hidden states to reduce errors from reusing another agent’s adapted cache. The authors report up to 3.1x faster TTFT, 2.3x better per-request throughput, and an average accuracy drop of 1.1 points for ReBaseShared versus no cache sharing. HF Daily Papers' note

score 4

Categories: Research