SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
The paper trains a VLA policy to absorb object-level 3D priors without needing 3D inputs at test time.
The method uses frozen SAM3D features as a training-time teacher for a policy built on $\pi_0$. It localizes task-relevant objects, masks them, extracts dense 3D representations, and aligns those with the model’s intermediate visual features. The authors report 99.1% on LIBERO and an average length of 4.11 on CALVIN, with real-world gains in long-horizon tasks involving changing target objects. ArXiv · AI/CL/LG's note
The method uses frozen SAM3D features as a training-time teacher for a policy built on $\pi_0$. It localizes task-relevant objects, masks them, extracts dense 3D representations, and aligns those with the model’s intermediate visual features. The authors report 99.1% on LIBERO and an average length of 4.11 on CALVIN, with real-world gains in long-horizon tasks involving changing target objects. ArXiv · AI/CL/LG's note
score 5