Consilience for Verifier-Free Test-Time Scaling
The paper argues that confidence-only selection can reward answers that never really explored the problem.
The authors say verifier-free test-time scaling breaks down on complex tasks when rollouts are ranked only by confidence. Their proposed “consilience” metric looks for low confidence early in reasoning and high confidence at the end. It penalizes rollouts that start confidently and demands final certainty. They report gains over existing baselines on graduate-level math and free-form code generation. ArXiv · AI/CL/LG's note
The authors say verifier-free test-time scaling breaks down on complex tasks when rollouts are ranked only by confidence. Their proposed “consilience” metric looks for low confidence early in reasoning and high confidence at the end. It penalizes rollouts that start confidently and demands final certainty. They report gains over existing baselines on graduate-level math and free-form code generation. ArXiv · AI/CL/LG's note
score 5