Megadose AI progress, ranked and analyzed.

AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

· ArXiv · AI/CL/LG ·
AdaFlash targets a variance problem in diffusion-based speculative decoding and reports up to about 66% higher throughput than prior state of the art.

The paper says diffusion drafters gain speed from parallel denoising, but their bidirectional attention makes acceptance rates unstable across domains and draft quality uneven across token positions. AdaFlash combines on-policy distillation with reverse-KL training to reduce domain-level variance. It also adds an adaptive length head that changes candidate sequence length during decoding to cut target-model verification cost. ArXiv · AI/CL/LG's note

score 5

Categories: Research