Megadose AI progress, ranked and analyzed.

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference

· ArXiv · AI/CL/LG ·
PyroDash reports a token-level handoff scheme that can beat an LLM-only math baseline while using less of the large model.

The paper’s small model learns to emit a control token when it wants help, then sends the partial reasoning trace to a frozen LLM for one completion handoff. Its training combines control-token learning, offloading-focused fine-tuning, and cost-aware GRPO alignment. On five math reasoning benchmarks, one setting reached 64.04% average accuracy, 6.36 points above the LLM-only baseline, with 20.4% lower cost. A higher-cost-penalty setting cut total cost from $49.36 to $1.78 while using the LLM on only 1.90% of tokens. ArXiv · AI/CL/LG's note

score 4

Categories: Research