VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
VeriFine makes the verifier evolve alongside the embodied agent it is judging.
The paper describes an agent harness where the policy, curriculum, and judge improve together as new failure modes appear. A rubric judge identifies recurring failures, builds adaptive training data, and guides policy optimization. When verification stalls, the system asks humans to help resolve informative failure cases and recalibrates the judge. Experiments on driving and robot navigation tasks report continued gains for both the policy and the judge across reinforcement learning and supervised fine-tuning. HF Daily Papers' note
The paper describes an agent harness where the policy, curriculum, and judge improve together as new failure modes appear. A rubric judge identifies recurring failures, builds adaptive training data, and guides policy optimization. When verification stalls, the system asks humans to help resolve informative failure cases and recalibrates the judge. Experiments on driving and robot navigation tasks report continued gains for both the policy and the judge across reinforcement learning and supervised fine-tuning. HF Daily Papers' note
score 5