Megadose AI progress, ranked and analyzed.

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

· ArXiv · AI/CL/LG ·
Training models to estimate confidence mid-reasoning cut token use by up to 25% without adding a stopping rule.

The paper fine-tunes reasoning models on 600 problems to predict confidence at intermediate points in their own reasoning traces. That confidence signal is only used during training, with no loss term for shorter outputs and no early-stopping mechanism at inference. Across Gemma, Qwen, Nemotron, and GPT-OSS models, the authors report matched-accuracy reductions in generated tokens on math, science, and coding benchmarks. Their analysis says the training mostly preserves the models’ broader reasoning patterns rather than pruning specific behaviors.

ArXiv · AI/CL/LG's note

score 5

Categories: Research