Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning
The paper argues that token-level reward signals in long-CoT training are being assigned too evenly, and often to the wrong tokens.
The authors test privileged self-teacher shifts by rescoring fixed reasoning traces under correct and incorrect outcome conditions, finding that many tokens move the same way in both cases. They say large shifts tend to land on substitutable surface-form tokens rather than the reasoning tokens that matter. Their proposed CSCR method downweights highly sensitive tokens while preserving the verifier’s overall reward direction. On long-CoT math benchmarks, CSCR beats a GRPO baseline with the same number of policy updates. HF Daily Papers' note
The authors test privileged self-teacher shifts by rescoring fixed reasoning traces under correct and incorrect outcome conditions, finding that many tokens move the same way in both cases. They say large shifts tend to land on substitutable surface-form tokens rather than the reasoning tokens that matter. Their proposed CSCR method downweights highly sensitive tokens while preserving the verifier’s overall reward direction. On long-CoT math benchmarks, CSCR beats a GRPO baseline with the same number of policy updates. HF Daily Papers' note
score 5