Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
The paper claims OPPD lets a model absorb power-sampling gains into a single generation.
The method trains on candidates generated by the same model, weighted by a frozen teacher’s sharpened sequence-level distribution. Reported gains reach 23.0 points on MATH500 and 27.3 on GSM8K over the untrained model at the same temperature. One generation also beats published 64-candidate power sampling by 2.4 and 3.5 points on those benchmarks. The authors say it uses no reference answers, complements GRPO, and transfers from math training to HumanEval. HF Daily Papers' note
The method trains on candidates generated by the same model, weighted by a frozen teacher’s sharpened sequence-level distribution. Reported gains reach 23.0 points on MATH500 and 27.3 on GSM8K over the untrained model at the same temperature. One generation also beats published 64-candidate power sampling by 2.4 and 3.5 points on those benchmarks. The authors say it uses no reference answers, complements GRPO, and transfers from math training to HumanEval. HF Daily Papers' note
score 4