SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
The paper’s claim is that fixed RLVR-plus-OPD training can collapse entropy, and SAF avoids it by reshaping only the OPD advantage.
The authors say OPD’s token-level advantages can overwhelm RLVR’s bounded signal, while sustained teacher pressure can limit exploration. SAF adds sparsify-then-compress magnitude control and warm-up-then-anneal timing control, with each stage independently switchable. Tested with GRPO on seven math reasoning and code generation benchmarks using Qwen3-1.7B/4B/8B, it beats fixed-coefficient fusion by 0.51–2.70% across six model-domain settings. The paper is marked “working in progress.” HF Daily Papers' note
The authors say OPD’s token-level advantages can overwhelm RLVR’s bounded signal, while sustained teacher pressure can limit exploration. SAF adds sparsify-then-compress magnitude control and warm-up-then-anneal timing control, with each stage independently switchable. Tested with GRPO on seven math reasoning and code generation benchmarks using Qwen3-1.7B/4B/8B, it beats fixed-coefficient fusion by 0.51–2.70% across six model-domain settings. The paper is marked “working in progress.” HF Daily Papers' note
score 5