On-Policy Delta Distillation
The paper’s key move is to distill the teacher’s reasoning gain, not its whole output distribution.
The authors define a “delta signal” as the difference between a reasoning-tuned teacher and its base model before instruction tuning. That signal is used as the reward for on-policy distillation, aiming to transfer the tuning-induced reasoning behavior more directly. They report consistent gains over conventional on-policy distillation on math, science, and code-reasoning benchmarks, with strong results after a short post-training period. ArXiv · AI/CL/LG's note
The authors define a “delta signal” as the difference between a reasoning-tuned teacher and its base model before instruction tuning. That signal is used as the reward for on-policy distillation, aiming to transfer the tuning-induced reasoning behavior more directly. They report consistent gains over conventional on-policy distillation on math, science, and code-reasoning benchmarks, with strong results after a short post-training period. ArXiv · AI/CL/LG's note
score 5