Megadose AI progress, ranked daily.

Best Practice Critic Optimization

· HF Daily Papers ·
The paper argues that a carefully constrained critic can replace multi-sample group advantage methods while using one response per prompt.

BPCO combines DPPO, bounded value predictions, Monte Carlo targets, unnormalized advantages, and length-adaptive GAE to stabilize critic-based training. The critic can see reward-defining information, such as reference answers or rubrics, during training while the policy cannot. In math reasoning experiments from 1.5B models to 30B-A3B MoE models, it consistently beats a strong critic baseline and matches or exceeds a group-based baseline. The same recipe also improves rubric-based reward learning. HF Daily Papers' note

score 5

Categories: Research