Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
The paper’s fix is to make vision pass through a pose-supervised bottleneck before it can drive robot actions.
Latent Interface Training first teaches an action model from language, robot state, and terminal end-effector pose, without images. It then adds a visual interface supervised to recover that pose, limiting image conditioning to task-relevant spatial information. The authors report 3.87 to 10.70 point gains on LIBERO-Plus across four model families, while maintaining or improving average LIBERO success. Real-world tests showed 13.30 to 16.70 point gains across three tasks under new cameras, lighting changes, and distractors. HF Daily Papers' note
Latent Interface Training first teaches an action model from language, robot state, and terminal end-effector pose, without images. It then adds a visual interface supervised to recover that pose, limiting image conditioning to task-relevant spatial information. The authors report 3.87 to 10.70 point gains on LIBERO-Plus across four model families, while maintaining or improving average LIBERO success. Real-world tests showed 13.30 to 16.70 point gains across three tasks under new cameras, lighting changes, and distractors. HF Daily Papers' note
score 5