WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
Wavefront Decoding reports up to 4.81x faster decoding for looped language models without training.
The paper targets looped LMs, where repeated calls to a shared recurrent block slow generation. Its method batches shallow draft states and deeper verification states together along a “wavefront,” instead of separating draft and verify phases. Reported speedups are 2.42x on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.5B with cross-recurrence KV sharing. HF Daily Papers' note
The paper targets looped LMs, where repeated calls to a shared recurrent block slow generation. Its method batches shallow draft states and deeper verification states together along a “wavefront,” instead of separating draft and verify phases. Reported speedups are 2.42x on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, rising to 4.81x on Huginn-3.5B with cross-recurrence KV sharing. HF Daily Papers' note
score 4