ECHO: Early-layer Collaborative Hierarchical Orchestration with Bonus Logits in Speculative Decoding
ECHO claims faster draft-model-free speculative decoding by splitting cheap early-layer drafting from full-model verification.
The paper says early-layer “bonus logits” explore draft trees in a high-frequency inner loop, while final layers verify and correct candidates in a lower-frequency outer loop. The authors report 2.4x to 2.9x speedups across benchmarks, with higher mean accepted tokens than cited state-of-the-art baselines. They say it adds no deployment parameters and has negligible engineering overhead, but optimal acceleration depends on one-shot fine-tuning. ArXiv · AI/CL/LG's note
The paper says early-layer “bonus logits” explore draft trees in a high-frequency inner loop, while final layers verify and correct candidates in a lower-frequency outer loop. The authors report 2.4x to 2.9x speedups across benchmarks, with higher mean accepted tokens than cited state-of-the-art baselines. They say it adds no deployment parameters and has negligible engineering overhead, but optimal acceleration depends on one-shot fine-tuning. ArXiv · AI/CL/LG's note
score 5