Megadose Built for builders and researchers.

Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

· ArXiv · AI/CL/LG ·
A frozen behavior-cloning robot policy improved by training only a small Q-function on its own rollouts.

The paper proposes Q-Planning, which uses an off-policy Q-function to guide action selection from a large visuomotor imitation policy without changing the imitation model’s weights. In tests, ten self-improvement iterations raised LIBERO-10 from 93% to 99% and RoboTwin from 83.8% to 91.4%. On real bimanual tasks, stack-cups improved from 40% to 90% and insert-wallet from 25% to 80% in five iterations, with no added human demonstrations. The authors say it was the only compared method to improve stably from failures under the same online budget. ArXiv · AI/CL/LG's note

score 6

Categories: Research