Semifactual Credit-Augmented Policy Optimization
SCAPO tries to stop RL training from rewarding tokens that change under irrelevant prompt variations.
The paper argues that GRPO can reinforce spurious prompt dependence because it gives every response token the same outcome-derived advantage. SCAPO instead measures how token probabilities drift under semifactual prompt interventions and downweights unstable tokens early in training. On Qwen3-4B-Base and Qwen3-1.7B-Base, it reports AIME 2024-2026 gains over GRPO of 5.63 and 4.17 percentage points. HF Daily Papers' note
The paper argues that GRPO can reinforce spurious prompt dependence because it gives every response token the same outcome-derived advantage. SCAPO instead measures how token probabilities drift under semifactual prompt interventions and downweights unstable tokens early in training. On Qwen3-4B-Base and Qwen3-1.7B-Base, it reports AIME 2024-2026 gains over GRPO of 5.63 and 4.17 percentage points. HF Daily Papers' note
score 4