DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation
DuoOPD changes distillation feedback based on whether the teacher, the student, both, or neither got the task right.
The paper says standard on-policy distillation can penalize a student’s correct answer when the teacher fails. DuoOPD instead uses the student outcome to set feedback direction, then uses the joint teacher-student result to decide how teacher information is applied. In tests across Qwen3 and Llama, it beat five baselines in mean macro accuracy, including OPD by 2.58 and 5.98 points. Ablations credited most of the gain to the joint-outcome design, not just correctness gating. HF Daily Papers' note
The paper says standard on-policy distillation can penalize a student’s correct answer when the teacher fails. DuoOPD instead uses the student outcome to set feedback direction, then uses the joint teacher-student result to decide how teacher information is applied. In tests across Qwen3 and Llama, it beat five baselines in mean macro accuracy, including OPD by 2.58 and 5.98 points. Ablations credited most of the gain to the joint-outcome design, not just correctness gating. HF Daily Papers' note
score 4