Megadose AI progress, ranked daily.

AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs

· ArXiv · AI/CL/LG ·
AsymSpec lets a small drafter see the full context while the large verifier works from a compressed one.

The paper targets agentic LLM pipelines where accumulated retrieval, tool, and chat context raises inference cost. Its method uses contrastive logit fusion and a divergence-aware acceptance gate to keep speculative decoding stable under mismatched context. In the reported evaluations, it averages about 90% of full-context accuracy, with 1.3-1.7x throughput gains and 0.2-0.3x compute cost on isolated text capabilities. ArXiv · AI/CL/LG's note

score 5

Categories: Research