Megadose Built for builders and researchers.

Semifactual Credit-Augmented Policy Optimization

· HF Daily Papers ·
SCAPO tries to stop RL training from rewarding tokens that change under irrelevant prompt variations.

The paper argues that GRPO can reinforce spurious prompt dependence because it gives every response token the same outcome-derived advantage. SCAPO instead measures how token probabilities drift under semifactual prompt interventions and downweights unstable tokens early in training. On Qwen3-4B-Base and Qwen3-1.7B-Base, it reports AIME 2024-2026 gains over GRPO of 5.63 and 4.17 percentage points. HF Daily Papers' note

score 4

Categories: Research