How to Train a Critic Stably and Efficiently
BPCO trains a critic for LLM reinforcement learning that can match group-based methods while using one sampled response per prompt.
The paper targets the instability that has kept critic-based training behind methods like GRPO. Its recipe combines DPPO, bounded value predictions, Monte Carlo targets, unnormalized policy advantages, and length-adaptive GAE. The critic can also see reward-defining information, such as a reference answer or rubric, because it is used only during training. In math-reasoning experiments across 1.5B to 30B-A3B models, BPCO consistently beats a strong critic baseline and matches or exceeds a group-based baseline. ArXiv · AI/CL/LG's note
The paper targets the instability that has kept critic-based training behind methods like GRPO. Its recipe combines DPPO, bounded value predictions, Monte Carlo targets, unnormalized policy advantages, and length-adaptive GAE. The critic can also see reward-defining information, such as a reference answer or rubric, because it is used only during training. In math-reasoning experiments across 1.5B to 30B-A3B models, BPCO consistently beats a strong critic baseline and matches or exceeds a group-based baseline. ArXiv · AI/CL/LG's note
score 5