Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models
Dr. OPD weights a teacher model’s token-level guidance instead of treating every token as equally useful.
The paper frames that weighting as a bilevel optimization problem, with token weights chosen to improve the student model’s expected reward. Its solver alternates between closed-form weight updates and one gradient step on the weighted distillation objective. The authors report gains on math and code distillation tests, including a 9.7-point average math improvement over vanilla OPD in strong-to-weak distillation. In that setting, they say the smaller student can surpass its larger teacher. ArXiv · AI/CL/LG's note
The paper frames that weighting as a bilevel optimization problem, with token weights chosen to improve the student model’s expected reward. Its solver alternates between closed-form weight updates and one gradient step on the weighted distillation objective. The authors report gains on math and code distillation tests, including a 9.7-point average math improvement over vanilla OPD in strong-to-weak distillation. In that setting, they say the smaller student can surpass its larger teacher. ArXiv · AI/CL/LG's note
score 5