Latent On-Policy Self-Distillation
The paper’s core move is to make the teacher’s privileged context learnable, instead of hand-designing it.
LOPD retrieves relevant experience and turns it into continuous latent tokens that condition a self-teacher. The student acts on its own task and interaction history, then gets dense token-level supervision along the visited trajectory. The authors report better results than RLVR and several OPSD baselines on agentic tool use and code generation, with less than 30% of the rollout budget used by GRPO and Skill-SD. Ablations are presented as evidence that the learnable privileged context is driving the gains. HF Daily Papers' note
LOPD retrieves relevant experience and turns it into continuous latent tokens that condition a self-teacher. The student acts on its own task and interaction history, then gets dense token-level supervision along the visited trajectory. The authors report better results than RLVR and several OPSD baselines on agentic tool use and code generation, with less than 30% of the rollout budget used by GRPO and Skill-SD. Ablations are presented as evidence that the learnable privileged context is driving the gains. HF Daily Papers' note
score 4