AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
AgSpec speeds coding-agent generation by making retrieval drafts match what the agent is actually likely to reuse.
The paper says existing retrieval-based speculative decoding misses useful agent context or stores it in the wrong format. AgSpec builds draft sources from the current session, workspace, and global corpora, including opened files in the agent’s emission format. It also caps draft length per agent from offline profiling and adjusts online using verification feedback. On two repository-level multi-agent coding benchmarks, it reports throughput gains up to 4.37x at batch size 1 and 4.76x at batch size 16 over autoregressive decoding. HF Daily Papers' note
The paper says existing retrieval-based speculative decoding misses useful agent context or stores it in the wrong format. AgSpec builds draft sources from the current session, workspace, and global corpora, including opened files in the agent’s emission format. It also caps draft length per agent from offline profiling and adjusts online using verification feedback. On two repository-level multi-agent coding benchmarks, it reports throughput gains up to 4.37x at batch size 1 and 4.76x at batch size 16 over autoregressive decoding. HF Daily Papers' note
score 4