Megadose AI progress, ranked and analyzed.

RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer

· ArXiv · AI/CL/LG ·
The paper targets the exact reasoning “pivot” tokens that matter most for transferring math reasoning across languages.

RP-OPSD uses differences between teacher views with and without an English reference solution to identify where privileged distillation should focus. The authors test it on math reasoning benchmarks across 17 languages and multiple difficulty levels. They report gains over multilingual reasoning baselines and other OPSD variants. Their analysis says the method emphasizes reasoning-control and state-update tokens while reducing weight on surface wording. ArXiv · AI/CL/LG's note

score 4

Categories: Research