Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
The paper’s claimed gain is making variable-depth looped LM inference batch efficiently enough to approach its theoretical speedup.
The authors introduce continuous depth batching, which rebuilds batches between loop steps so tokens that need different numbers of recurrent passes can still run together. The system also handles looped KV caching and predicts token exits early to prepare later batches asynchronously. Tests on Ouro 1.4B and Huginn 3.5B show fully looped designs benefit most, while large non-looped components make scheduling harder. CDB reaches up to 99% of the estimated maximum available speedup in their experiments. HF Daily Papers' note
The authors introduce continuous depth batching, which rebuilds batches between loop steps so tokens that need different numbers of recurrent passes can still run together. The system also handles looped KV caching and predicts token exits early to prepare later batches asynchronously. Tests on Ouro 1.4B and Huginn 3.5B show fully looped designs benefit most, while large non-looped components make scheduling harder. CDB reaches up to 99% of the estimated maximum available speedup in their experiments. HF Daily Papers' note
score 5