Enhancing Rubric-based RL via Self-Distillation
CriPO targets rubric signals that standard RL either never sees or accidentally suppresses.
The paper identifies two failure modes in rubric-based RL: unexplored criteria and “suppressed criteria,” where useful satisfied criteria get lost after scalar reward aggregation. It says suppressed criteria appear in more than 57% of samples during training, averaging 1.8 per sample. CriPO uses on-policy self-distillation to inject missing behaviors for unexplored criteria and preserve criterion-relevant tokens from negative-advantage rollouts. In medicine and science benchmarks, the authors report stronger final performance with about 2x fewer optimization steps. ArXiv · AI/CL/LG's note
The paper identifies two failure modes in rubric-based RL: unexplored criteria and “suppressed criteria,” where useful satisfied criteria get lost after scalar reward aggregation. It says suppressed criteria appear in more than 57% of samples during training, averaging 1.8 per sample. CriPO uses on-policy self-distillation to inject missing behaviors for unexplored criteria and preserve criterion-relevant tokens from negative-advantage rollouts. In medicine and science benchmarks, the authors report stronger final performance with about 2x fewer optimization steps. ArXiv · AI/CL/LG's note
score 5