Megadose AI progress, ranked and analyzed.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

· HF Daily Papers ·
CPO replaces entropy with token-level disagreement as the correctness signal for RLVR advantage shaping.

The paper argues entropy cannot tell useful uncertainty from confusion. Its Contrastive Policy Optimization compares reference-guided and vanilla generation distributions to shape advantages around token-level correctness. The authors say the method resolves the zero-advantage problem and outperforms entropy-based RLVR on in-domain and out-of-domain benchmarks. HF Daily Papers' note

score 5

Categories: Research