Semifactual Credit-Augmented Policy Optimization
SCAPO changes RLVR credit assignment so unstable tokens get less reinforcement early in training.
The paper tests “semifactual” prompt changes that keep the same problem and answer, then measures which token choices drift under those changes. It argues GRPO can over-credit every token in a successful response, including tokens tied to irrelevant prompt features. SCAPO uses token-level stability scores to dampen credit for less stable tokens, without rewarding stability by itself. On Qwen3-4B-Base and Qwen3-1.7B-Base, it reports AIME 2024-2026 gains over GRPO of 5.63 and 4.17 points.
ArXiv · AI/CL/LG's note
The paper tests “semifactual” prompt changes that keep the same problem and answer, then measures which token choices drift under those changes. It argues GRPO can over-credit every token in a successful response, including tokens tied to irrelevant prompt features. SCAPO uses token-level stability scores to dampen credit for less stable tokens, without rewarding stability by itself. On Qwen3-4B-Base and Qwen3-1.7B-Base, it reports AIME 2024-2026 gains over GRPO of 5.63 and 4.17 points.
ArXiv · AI/CL/LG's note
score 4