Megadose AI progress, ranked and analyzed.

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

· ArXiv · AI/CL/LG ·
The paper argues that test-time scaling results are not comparable unless the full inference protocol is reported.

It separates reasoning-time compute into single-trajectory deliberation, candidate sampling with voting or verification, and search over partial prefixes. The authors treat the evaluated object as the whole inference system, not just the base model or a bank of sampled answers. They call for protocol-matched compute accounting, uncertainty reporting, and reproducibility artifacts for exact replay or distributional reproduction. The paper also says it is releasing more than 2 billion full reasoning traces with added verifier and token-level signals. ArXiv · AI/CL/LG's note

score 5

Categories: Research