KV-streams for Efficient Compaction in Agentic Reinforcement Learning
KV-streams keep the KV cache moving through compaction, cutting agentic RL training time without reported performance loss.
The paper says repeated context prefills are a throughput bottleneck when training long-horizon agentic LLMs. Its proposed plug-and-play method streams the KV cache forward instead of flushing it after each compaction step. Across three compaction strategies, the authors report 2.6x to 5x wall-clock training speedups. They also find the streamed cache can behave like recurrent state, preserving information no longer present in the visible context. ArXiv · AI/CL/LG's note
The paper says repeated context prefills are a throughput bottleneck when training long-horizon agentic LLMs. Its proposed plug-and-play method streams the KV cache forward instead of flushing it after each compaction step. Across three compaction strategies, the authors report 2.6x to 5x wall-clock training speedups. They also find the streamed cache can behave like recurrent state, preserving information no longer present in the visible context. ArXiv · AI/CL/LG's note
score 5