Receiver-Conditioned Latent Communication gives 94% CacheBack
CacheBack filters shared agent state around what the receiver says it needs, cutting transferred KV cache by 75% in the paper’s FanOutQA test.
The authors frame the problem as latent communication becoming too large as context and agent count grow. Their method sends a small receiver-side description back to the sender, then uses attention weights to compress the sender’s KV cache without training. With Qwen 3 on FanOutQA, they report a 14.7-point accuracy gain and 3.2x lower median completion latency versus text communication. They also report similar gains across dense Transformers, Mamba-attention hybrids, and sliding-window attention models. HF Daily Papers' note
The authors frame the problem as latent communication becoming too large as context and agent count grow. Their method sends a small receiver-side description back to the sender, then uses attention weights to compress the sender’s KV cache without training. With Qwen 3 on FanOutQA, they report a 14.7-point accuracy gain and 3.2x lower median completion latency versus text communication. They also report similar gains across dense Transformers, Mamba-attention hybrids, and sliding-window attention models. HF Daily Papers' note
score 5