SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models
SpecFold reuses near-duplicate branch computation inside DLLM speculative verification to cut decoding cost.
The paper says draft branches mostly inherit parent tokens, leaving many hidden states similar across branches. SpecFold adds token-level residual gating, folded attention, and folded FFN reuse, implemented with Triton for sparse multi-branch execution. Across two DLLM families, five models, and five benchmarks, it reports up to 1.64x throughput over Spiffy and 1.99x over vanilla decoding with comparable task performance. Source: HF Daily Papers' note.
The paper says draft branches mostly inherit parent tokens, leaving many hidden states similar across branches. SpecFold adds token-level residual gating, folded attention, and folded FFN reuse, implemented with Triton for sparse multi-branch execution. Across two DLLM families, five models, and five benchmarks, it reports up to 1.64x throughput over Spiffy and 1.99x over vanilla decoding with comparable task performance. Source: HF Daily Papers' note.
score 5