Best Practice Critic Optimization
The paper argues that a carefully constrained critic can replace multi-sample group advantage methods while using one response per prompt.
BPCO combines DPPO, bounded value predictions, Monte Carlo targets, unnormalized advantages, and length-adaptive GAE to stabilize critic-based training. The critic can see reward-defining information, such as reference answers or rubrics, during training while the policy cannot. In math reasoning experiments from 1.5B models to 30B-A3B MoE models, it consistently beats a strong critic baseline and matches or exceeds a group-based baseline. The same recipe also improves rubric-based reward learning. HF Daily Papers' note
BPCO combines DPPO, bounded value predictions, Monte Carlo targets, unnormalized advantages, and length-adaptive GAE to stabilize critic-based training. The critic can see reward-defining information, such as reference answers or rubrics, during training while the policy cannot. In math reasoning experiments from 1.5B models to 30B-A3B MoE models, it consistently beats a strong critic baseline and matches or exceeds a group-based baseline. The same recipe also improves rubric-based reward learning. HF Daily Papers' note
score 5