Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
CPO replaces entropy with token-level disagreement as the correctness signal for RLVR advantage shaping.
The paper argues entropy cannot tell useful uncertainty from confusion. Its Contrastive Policy Optimization compares reference-guided and vanilla generation distributions to shape advantages around token-level correctness. The authors say the method resolves the zero-advantage problem and outperforms entropy-based RLVR on in-domain and out-of-domain benchmarks. HF Daily Papers' note
The paper argues entropy cannot tell useful uncertainty from confusion. Its Contrastive Policy Optimization compares reference-guided and vanilla generation distributions to shape advantages around token-level correctness. The authors say the method resolves the zero-advantage problem and outperforms entropy-based RLVR on in-domain and out-of-domain benchmarks. HF Daily Papers' note
score 5