CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing
CoBa reports near-best-of-16 accuracy while spending far fewer parameter-weighted tokens.
The paper frames test-time reasoning as a routing problem: spend the next compute unit on generation, verification, or stop. CoBa starts with a small candidate set, uses cheap verification widely, then sends uncertain or valuable cases to stronger checks. Across 3,129 evaluations, CoBa-Routed-Strong reached 85.13% macro accuracy, close to a self-evaluation weighted-voting proxy at 85.20%, with 49.1% fewer parameter-weighted tokens. It also matched best-of-16 majority voting within 0.01 macro-accuracy points while using 58.9% fewer parameter-weighted tokens, though paired tests still show a small best-of-16 edge at higher cost. ArXiv · AI/CL/LG's note
The paper frames test-time reasoning as a routing problem: spend the next compute unit on generation, verification, or stop. CoBa starts with a small candidate set, uses cheap verification widely, then sends uncertain or valuable cases to stronger checks. Across 3,129 evaluations, CoBa-Routed-Strong reached 85.13% macro accuracy, close to a self-evaluation weighted-voting proxy at 85.20%, with 49.1% fewer parameter-weighted tokens. It also matched best-of-16 majority voting within 0.01 macro-accuracy points while using 58.9% fewer parameter-weighted tokens, though paired tests still show a small best-of-16 edge at higher cost. ArXiv · AI/CL/LG's note
score 5