Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
HeteroFold transfers KV caches between different LLM families without making the receiving model redo prefill.
The paper targets multi-agent systems where one model has already processed a long shared context and another model would normally have to process it again. HeteroFold keeps both models frozen, then aligns their structures, maps the sender cache into the receiver’s space, and calibrates it to preserve receiver behavior. In the reported tests, it led all cache-transfer methods across four long-context benchmarks and most short-context settings, while matching text-based communication on the multi-agent benchmark. For Llama-3.1-8B to Ministral-3-14B at 32K context, the authors report a 10.7x speedup over Native Prefill. HF Daily Papers' note
The paper targets multi-agent systems where one model has already processed a long shared context and another model would normally have to process it again. HeteroFold keeps both models frozen, then aligns their structures, maps the sender cache into the receiver’s space, and calibrates it to preserve receiver behavior. In the reported tests, it led all cache-transfer methods across four long-context benchmarks and most short-context settings, while matching text-based communication on the multi-agent benchmark. For Llama-3.1-8B to Ministral-3-14B at 32K context, the authors report a 10.7x speedup over Native Prefill. HF Daily Papers' note
score 5