Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves
The paper argues that test-time scaling curves need simultaneous certification across budgets, not point-by-point confidence.
A 100-question benchmark can require 192,000 generated answers to certify 64 budgets within ±1/32 at 95% under a fixed exact-binomial design. The authors say much of that cost is unnecessary because benchmark variance is mostly between questions, which an audit revisiting every question does not need to repay. Their paired audit used fewer answers than competing certified audits on held-out score pools, and certified a new MMLU-Pro curve with 79,133 answers. ArXiv · AI/CL/LG's note
A 100-question benchmark can require 192,000 generated answers to certify 64 budgets within ±1/32 at 95% under a fixed exact-binomial design. The authors say much of that cost is unnecessary because benchmark variance is mostly between questions, which an audit revisiting every question does not need to repay. Their paired audit used fewer answers than competing certified audits on held-out score pools, and certified a new MMLU-Pro curve with 79,133 answers. ArXiv · AI/CL/LG's note
score 5