SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
SimpleOPD lets short-context models learn long-context proof reasoning despite tokenizer mismatch.
The paper distills reasoning from SU-01 into shorter-context student models by aligning shared text spans rather than requiring matching tokenizers. It adds a student reference KL loss and masks termination-token advantages to control runaway response length and training instability. The authors report consistent math-reasoning gains across Qwen, Intern-S2, GLM, and Gemma students, with Intern-S2-Preview rising 21.2 points on ProofBench to 55.2. They also report gains on HLE and HiPhO, arguing the transferred reasoning generalizes beyond math proof training. HF Daily Papers' note
The paper distills reasoning from SU-01 into shorter-context student models by aligning shared text spans rather than requiring matching tokenizers. It adds a student reference KL loss and masks termination-token advantages to control runaway response length and training instability. The authors report consistent math-reasoning gains across Qwen, Intern-S2, GLM, and Gemma students, with Intern-S2-Preview rising 21.2 points on ProofBench to 55.2. They also report gains on HLE and HiPhO, arguing the transferred reasoning generalizes beyond math proof training. HF Daily Papers' note
score 5