An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
LSPD reframes on-policy distillation as KL-regularized policy optimization, then uses RL-style exploration and off-policy reuse to cut rollout needs.
The paper connects OPD’s reverse-KL objective to reinforcement learning and proposes Least-Square Policy Distillation as the resulting framework. LSPD is reported to beat existing distillation baselines across six math reasoning benchmarks, with an average +1.59 Avg@16 gain. Its Pass@k results up to k=64 suggest it preserves more policy diversity as more samples are drawn. A fully off-policy variant matched vanilla OPD while using only the first 25% of rollout batches. Source: HF Daily Papers' note.
The paper connects OPD’s reverse-KL objective to reinforcement learning and proposes Least-Square Policy Distillation as the resulting framework. LSPD is reported to beat existing distillation baselines across six math reasoning benchmarks, with an average +1.59 Avg@16 gain. Its Pass@k results up to k=64 suggest it preserves more policy diversity as more samples are drawn. A fully off-policy variant matched vanilla OPD while using only the first 25% of rollout batches. Source: HF Daily Papers' note.
score 5