Megadose AI progress, ranked and analyzed.

On-Policy Delta Distillation

· ArXiv · AI/CL/LG ·
The paper’s key move is to distill the teacher’s reasoning gain, not its whole output distribution.

The authors define a “delta signal” as the difference between a reasoning-tuned teacher and its base model before instruction tuning. That signal is used as the reward for on-policy distillation, aiming to transfer the tuning-induced reasoning behavior more directly. They report consistent gains over conventional on-policy distillation on math, science, and code-reasoning benchmarks, with strong results after a short post-training period. ArXiv · AI/CL/LG's note

score 5

Categories: Research