Megadose AI progress, ranked and analyzed.

An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning

· HF Daily Papers ·
LSPD reframes on-policy distillation as KL-regularized policy optimization, then uses RL-style exploration and off-policy reuse to cut rollout needs.

The paper connects OPD’s reverse-KL objective to reinforcement learning and proposes Least-Square Policy Distillation as the resulting framework. LSPD is reported to beat existing distillation baselines across six math reasoning benchmarks, with an average +1.59 Avg@16 gain. Its Pass@k results up to k=64 suggest it preserves more policy diversity as more samples are drawn. A fully off-policy variant matched vanilla OPD while using only the first 25% of rollout batches. Source: HF Daily Papers' note.

score 5

Categories: Research