DLoop: Looped Speculative Decoding
DLoop cuts verification passes by letting confident draft models keep drafting before the target model checks the batch.
The paper says target models often accept every token from a draft stage, making some verification passes unnecessary. DLoop accumulates multiple draft stages when confidence stays high, then verifies the combined tokens at once. Its loop-aware training exposes the draft model to hidden states from unverified draft tokens so later loops stay reliable. Across several speculative decoding methods, the authors report 5% to 41% wall-clock speedup while preserving lossless decoding. HF Daily Papers' note
The paper says target models often accept every token from a draft stage, making some verification passes unnecessary. DLoop accumulates multiple draft stages when confidence stays high, then verifies the combined tokens at once. Its loop-aware training exposes the draft model to hidden states from unverified draft tokens so later loops stay reliable. Across several speculative decoding methods, the authors report 5% to 41% wall-clock speedup while preserving lossless decoding. HF Daily Papers' note
score 5