Megadose Built for builders and researchers.

How to Train a Critic Stably and Efficiently

· ArXiv · AI/CL/LG ·
BPCO trains a critic for LLM reinforcement learning that can match group-based methods while using one sampled response per prompt.

The paper targets the instability that has kept critic-based training behind methods like GRPO. Its recipe combines DPPO, bounded value predictions, Monte Carlo targets, unnormalized policy advantages, and length-adaptive GAE. The critic can also see reward-defining information, such as a reference answer or rubric, because it is used only during training. In math-reasoning experiments across 1.5B to 30B-A3B models, BPCO consistently beats a strong critic baseline and matches or exceeds a group-based baseline. ArXiv · AI/CL/LG's note

score 5

Categories: Research