Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
The paper argues that test-time scaling results are not comparable unless the full inference protocol is reported.
It separates reasoning-time compute into single-trajectory deliberation, candidate sampling with voting or verification, and search over partial prefixes. The authors treat the evaluated object as the whole inference system, not just the base model or a bank of sampled answers. They call for protocol-matched compute accounting, uncertainty reporting, and reproducibility artifacts for exact replay or distributional reproduction. The paper also says it is releasing more than 2 billion full reasoning traces with added verifier and token-level signals. ArXiv · AI/CL/LG's note
It separates reasoning-time compute into single-trajectory deliberation, candidate sampling with voting or verification, and search over partial prefixes. The authors treat the evaluated object as the whole inference system, not just the base model or a bank of sampled answers. They call for protocol-matched compute accounting, uncertainty reporting, and reproducibility artifacts for exact replay or distributional reproduction. The paper also says it is releasing more than 2 billion full reasoning traces with added verifier and token-level signals. ArXiv · AI/CL/LG's note
score 5