HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface
The paper proposes collecting policy-aware robot training data without running the policy on the robot.
HIL-UMI uses handheld Universal Manipulation Interface demonstrations while querying the current VLA model on the same observations. It saves new data when the human trajectory and policy prediction diverge, marking likely out-of-distribution states. A second loop identifies important low-advantage segments to refine a progress estimator, then uses that signal for advantage-conditioned behavioral cloning. In four real-world manipulation tasks, the authors report consistent gains over supervised fine-tuning and better Clean Up Table results than HG-DAgger with lower per-frame collection time.
ArXiv · AI/CL/LG's note
HIL-UMI uses handheld Universal Manipulation Interface demonstrations while querying the current VLA model on the same observations. It saves new data when the human trajectory and policy prediction diverge, marking likely out-of-distribution states. A second loop identifies important low-advantage segments to refine a progress estimator, then uses that signal for advantage-conditioned behavioral cloning. In four real-world manipulation tasks, the authors report consistent gains over supervised fine-tuning and better Clean Up Table results than HG-DAgger with lower per-frame collection time.
ArXiv · AI/CL/LG's note
score 5