$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning
The paper tests whether free-form language reasoning can improve long-horizon robot manipulation.
The authors introduce `R^3`, a post-training recipe for turning off-the-shelf vision-language models into robot reasoners. It uses expert reasoning traces first, then rubric-based reinforcement learning from offline action data. In Language Table and simulated bimanual grocery packing, the method improved exploration and generalization over instruction-only imitation baselines. ArXiv · AI/CL/LG's note
The authors introduce `R^3`, a post-training recipe for turning off-the-shelf vision-language models into robot reasoners. It uses expert reasoning traces first, then rubric-based reinforcement learning from offline action data. In Language Table and simulated bimanual grocery packing, the method improved exploration and generalization over instruction-only imitation baselines. ArXiv · AI/CL/LG's note
score 5