Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
The paper claims OPPD makes a model absorb power-sampling gains into a single generation.
Instead of drawing and scoring many candidates at inference time, the method trains on candidates from the model itself while a frozen teacher’s power distribution weights them. The authors report gains of up to 23.0 points on MATH500 and 27.3 on GSM8K over the untrained model at the same temperature. They also say one generation beats published 64-candidate power sampling by 2.4 and 3.5 points on those benchmarks. The method is trained only on math but is reported to improve HumanEval by up to 5.3 points. Source: ArXiv · AI/CL/LG's note
Instead of drawing and scoring many candidates at inference time, the method trains on candidates from the model itself while a frozen teacher’s power distribution weights them. The authors report gains of up to 23.0 points on MATH500 and 27.3 on GSM8K over the untrained model at the same temperature. They also say one generation beats published 64-candidate power sampling by 2.4 and 3.5 points on those benchmarks. The method is trained only on math but is reported to improve HumanEval by up to 5.3 points. Source: ArXiv · AI/CL/LG's note
score 6