Flux-OPD: On-Policy Distillation with Evolving Contexts
Flux-OPD uses changing context signals to stabilize on-policy distillation for open-ended LLM tasks.
The paper argues that fixed contexts lose supervisory value once absorbed by the student model. Its method anchors training on a context-free teacher, then injects context-conditioned differences as corrections. A conflict term is used to reduce the strength of corrections when the context-conditioned teachers disagree. The authors report that Flux-OPD beats existing OPD paradigms on open-ended tasks. HF Daily Papers' note
The paper argues that fixed contexts lose supervisory value once absorbed by the student model. Its method anchors training on a context-free teacher, then injects context-conditioned differences as corrections. A conflict term is used to reduce the strength of corrections when the context-conditioned teachers disagree. The authors report that Flux-OPD beats existing OPD paradigms on open-ended tasks. HF Daily Papers' note
score 4