Megadose AI progress, ranked and analyzed.

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

· ArXiv · AI/CL/LG ·
The paper trains a VLA policy to absorb object-level 3D priors without needing 3D inputs at test time.

The method uses frozen SAM3D features as a training-time teacher for a policy built on $\pi_0$. It localizes task-relevant objects, masks them, extracts dense 3D representations, and aligns those with the model’s intermediate visual features. The authors report 99.1% on LIBERO and an average length of 4.11 on CALVIN, with real-world gains in long-horizon tasks involving changing target objects. ArXiv · AI/CL/LG's note

score 5

Categories: Research