DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation
DiffGate adds teacher guidance only where student rollouts fail, using difficulty to control how much signal is applied.
The paper frames this as a bridge between on-policy distillation’s dense token guidance and GRPO’s outcome-level rewards. A verifier decides which failed trajectories get teacher supervision, while the teacher supplies token-level update directions inside those failures. In tests with Qwen3-0.6B and Qwen3-1.7B, DiffGate beat matched GRPO on code avg@8 and pass@8, and improved pass@8 across all four model-domain settings. HF Daily Papers' note
The paper frames this as a bridge between on-policy distillation’s dense token guidance and GRPO’s outcome-level rewards. A verifier decides which failed trajectories get teacher supervision, while the teacher supplies token-level update directions inside those failures. In tests with Qwen3-0.6B and Qwen3-1.7B, DiffGate beat matched GRPO on code avg@8 and pass@8, and improved pass@8 across all four model-domain settings. HF Daily Papers' note
score 5