SR-OPSD: Self-Referenced On-Policy Self-Distillation
The paper reframes OPSD’s moving self-teacher target as a controlled interpolation with a reference policy.
SR-OPSD separates the target placement from the projection used to train the student. It uses a token-level variational view to set the effective distillation target between the self-teacher and a reference policy, then applies Rényi divergence to tune projection geometry and sensitivity to density ratios. The authors report state-of-the-art or competitive results on scientific evaluation, math reasoning, and coding generation tasks across multiple LLMs. ArXiv · AI/CL/LG's note
SR-OPSD separates the target placement from the projection used to train the student. It uses a token-level variational view to set the effective distillation target between the self-teacher and a reference policy, then applies Rényi divergence to tune projection geometry and sensitivity to density ratios. The authors report state-of-the-art or competitive results on scientific evaluation, math reasoning, and coding generation tasks across multiple LLMs. ArXiv · AI/CL/LG's note
score 4