Megadose Built for builders and researchers.

Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds

· ArXiv · AI/CL/LG ·
The paper trains parallel speculative drafters against expected decoding rounds directly, instead of using block-local proxy losses.

Zhao and Cai model speculative decoding as a Markov reward process and define an Expected Decoding Rounds objective that equals the expected number of rounds. The objective weights rejection costs by state occupancy and uses no auxiliary hyperparameters, according to the abstract. They also derive an exact temporal-difference gradient for unbiased stochastic optimization from target-model rollouts. In tests, EDR finetuning improved DSpark and DFly on mean accepted length and beat existing objectives across nine math, code, and chat benchmarks. ArXiv · AI/CL/LG's note

score 5

Categories: Research