Megadose AI progress, ranked and analyzed.

xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding

· ArXiv · AI/CL/LG ·
xPress adds a causal refinement step to diffusion-based speculative decoding, improving draft acceptance without reverting to token-by-token drafting.

The paper says block-diffusion drafters can sample each draft position independently, producing tokens that look likely alone but fail conditional verification together. xPress refines the whole draft block in parallel so causal dependencies are restored across positions. On Qwen3-8B across seven math, code, and chat benchmarks, it reports about 30% higher average acceptance length and about 1.3x average end-to-end throughput versus dFlash. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research