Megadose AI progress, ranked and analyzed.

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

· ArXiv · AI/CL/LG ·
SwitchSD uses the target model’s own internal signals to decide when copying will actually help speculative decoding.

The paper says copy-based drafting can beat neural drafting in repetition-heavy text, but false copy triggers can hurt throughput. Its proposed lightweight probes detect “copy intent” in model representations with AUC above 0.99. SwitchSD then switches between EAGLE-style neural drafting and context-based copying. Across Llama and Qwen models, the authors report up to 15% throughput gains over baselines such as EAGLE3. ArXiv · AI/CL/LG's note

score 6

Categories: Research