Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
The paper says small models can close much of the reasoning gap at inference time by coordinating samplers at different sharpening levels.
Its method, Parallel Power Tempering, runs interacting replicas so some chains explore broader reasoning paths while sharper chains exploit higher-likelihood answers. The authors frame it as an alternative to RL post-training, avoiding parameter updates and external rewards. They also address truncation bias in earlier power samplers and test swap strategies under memory and compute limits. In their experiments, PPT improves single-chain power-sharpened sampling, beats RL-post-trained models, and reaches performance comparable to frontier models. HF Daily Papers' note
Its method, Parallel Power Tempering, runs interacting replicas so some chains explore broader reasoning paths while sharper chains exploit higher-likelihood answers. The authors frame it as an alternative to RL post-training, avoiding parameter updates and external rewards. They also address truncation bias in earlier power samplers and test swap strategies under memory and compute limits. In their experiments, PPT improves single-chain power-sharpened sampling, beats RL-post-trained models, and reaches performance comparable to frontier models. HF Daily Papers' note
score 4