More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It
Sharpening the “right” trajectories can still make reasoning results worse.
The paper says Power Sampling sometimes moves more probability mass onto correct paths while hurting downstream methods such as self-consistency. In their tests, accuracy fell by as much as 18.5 percentage points across models and reasoning benchmarks. The authors blame fixed-strength sharpening and the loss of broader reasoning-path support. Their proposed repair calibrates the deformation by problem and preserves more moderate-probability paths, improving same-budget weighted self-consistency. ArXiv · AI/CL/LG's note
The paper says Power Sampling sometimes moves more probability mass onto correct paths while hurting downstream methods such as self-consistency. In their tests, accuracy fell by as much as 18.5 percentage points across models and reasoning benchmarks. The authors blame fixed-strength sharpening and the loss of broader reasoning-path support. Their proposed repair calibrates the deformation by problem and preserves more moderate-probability paths, improving same-budget weighted self-consistency. ArXiv · AI/CL/LG's note
score 5