Megadose AI progress, ranked and analyzed.

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

· HF Daily Papers ·
SimpleOPD lets short-context models learn long-context proof reasoning despite tokenizer mismatch.

The paper distills reasoning from SU-01 into shorter-context student models by aligning shared text spans rather than requiring matching tokenizers. It adds a student reference KL loss and masks termination-token advantages to control runaway response length and training instability. The authors report consistent math-reasoning gains across Qwen, Intern-S2, GLM, and Gemma students, with Intern-S2-Preview rising 21.2 points on ProofBench to 55.2. They also report gains on HLE and HiPhO, arguing the transferred reasoning generalizes beyond math proof training. HF Daily Papers' note

score 5

Categories: Research