Generalizable Robotic Insertion with World Models
A single world-model policy handled unseen insertion parts at 56% zero-shot success, versus 7% for the model-free baseline.
The paper trains one model across up to 90 geometrically varied insertion tasks using robot proprioception and wrist-camera video. The authors say performance improves as more objects are added to training, suggesting the approach scales with broader task data. Fine-tuning on held-out objects was more data-efficient than training from scratch, and sometimes reached better final performance. The work is submitted as an IROS 2026 paper. ArXiv · AI/CL/LG's note
The paper trains one model across up to 90 geometrically varied insertion tasks using robot proprioception and wrist-camera video. The authors say performance improves as more objects are added to training, suggesting the approach scales with broader task data. Fine-tuning on held-out objects was more data-efficient than training from scratch, and sometimes reached better final performance. The work is submitted as an IROS 2026 paper. ArXiv · AI/CL/LG's note
score 5