Cliff: Learning Process Rewards from the First Mistake
Cliff trains from the first wrong step, not just the final answer.
The paper proposes using an off-the-shelf LLM teacher to mark where a rollout first fails. Tokens before that point get positive advantages; tokens after it get negative feedback. Across 12 scenarios, the authors report Cliff beating on-policy distillation by 15% and standard GRPO by 7%. Source: ArXiv · AI/CL/LG's note.
The paper proposes using an off-the-shelf LLM teacher to mark where a rollout first fails. Tokens before that point get positive advantages; tokens after it get negative feedback. Across 12 scenarios, the authors report Cliff beating on-policy distillation by 15% and standard GRPO by 7%. Source: ArXiv · AI/CL/LG's note.
score 5