DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees
DARTree reports lossless decoding speedups up to 9.73x by extending autoregressive correction from draft chains to draft trees.
The method is training-free and uses a pretrained AR correction head while building a fixed-width candidate tree in batches. It delays best-first pruning until after expansion, separating AR-head inference from sequential heap steps. Across seven math, code, and chat benchmarks, the authors say it produced the highest average acceptance length and speedup in all tested model-temperature settings. Source: ArXiv · AI/CL/LG's note.
The method is training-free and uses a pretrained AR correction head while building a fixed-width candidate tree in batches. It delays best-first pruning until after expansion, separating AR-head inference from sequential heap steps. Across seven math, code, and chat benchmarks, the authors say it produced the highest average acceptance length and speedup in all tested model-temperature settings. Source: ArXiv · AI/CL/LG's note.
score 5