Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
GC-OPD is meant to correct cases where a stronger teacher’s token guidance disagrees with task-level verifier rewards on long-context reasoning.
The paper says that, as inputs get longer, standard on-policy distillation scores become less aligned with verifier rewards on evidence-aggregation tasks. Its method normalizes verifier rewards and teacher OPD scores within rollout groups, then uses the residual disagreement to adjust token-level credit assignment without dropping the original OPD signal. On five long-context benchmarks, GC-OPD improves Qwen3-4B from 29.08 to 40.47 and Qwen3-8B from 35.12 to 44.65, slightly ahead of vanilla OPD under the same setup. ArXiv · AI/CL/LG's note
The paper says that, as inputs get longer, standard on-policy distillation scores become less aligned with verifier rewards on evidence-aggregation tasks. Its method normalizes verifier rewards and teacher OPD scores within rollout groups, then uses the residual disagreement to adjust token-level credit assignment without dropping the original OPD signal. On five long-context benchmarks, GC-OPD improves Qwen3-4B from 29.08 to 40.47 and Qwen3-8B from 35.12 to 44.65, slightly ahead of vanilla OPD under the same setup. ArXiv · AI/CL/LG's note
score 5