Megadose Built for builders and researchers.

Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

· HF Daily Papers ·
The paper claims OPPD lets a model absorb power-sampling gains into a single generation.

The method trains on candidates generated by the same model, weighted by a frozen teacher’s sharpened sequence-level distribution. Reported gains reach 23.0 points on MATH500 and 27.3 on GSM8K over the untrained model at the same temperature. One generation also beats published 64-candidate power sampling by 2.4 and 3.5 points on those benchmarks. The authors say it uses no reference answers, complements GRPO, and transfers from math training to HumanEval. HF Daily Papers' note

score 4

Categories: Research