Megadose AI progress, ranked and analyzed.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models

· ArXiv · AI/CL/LG ·
Lightning OPD 2.0 tries to separate useful teacher signal from recurring style differences in cross-teacher distillation.

The paper says OPD can stall when the SFT data and distillation supervision come from different teachers. Its method uses cross-fitted style residualization to estimate wording, formatting, and reasoning-cadence bias, then subtracts that component before the token-level update. On math and code benchmarks, it reports consistent gains over Lightning OPD in cross-teacher settings, including 82.4% on AIME 2024 and 63.0% on LiveCodeBench v5 from Klear-Reasoner-8B-SFT. Code is marked as coming soon. ArXiv · AI/CL/LG's note

score 5

Categories: Research