Megadose AI progress, ranked and analyzed.

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

· HF Daily Papers ·
The paper’s claim is that fixed RLVR-plus-OPD training can collapse entropy, and SAF avoids it by reshaping only the OPD advantage.

The authors say OPD’s token-level advantages can overwhelm RLVR’s bounded signal, while sustained teacher pressure can limit exploration. SAF adds sparsify-then-compress magnitude control and warm-up-then-anneal timing control, with each stage independently switchable. Tested with GRPO on seven math reasoning and code generation benchmarks using Qwen3-1.7B/4B/8B, it beats fixed-coefficient fusion by 0.51–2.70% across six model-domain settings. The paper is marked “working in progress.” HF Daily Papers' note

score 5

Categories: Research