See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
The paper argues that VLA models generalize better when scene geometry is expressed in the robot’s own coordinate frame.
The authors introduce robot-centric pointmaps: image-like inputs whose pixels store 3D scene coordinates in the robot frame. The format keeps the dense grid expected by pretrained 2D VLA models, so it can be added with minimal architecture changes. On RoboCasa, the method improves pi0.5 and SmolVLA and beats camera-viewpoint and 3D-aware baselines. In real-robot tests, the gain over RGB-only control is larger when the camera is moved to an unseen placement. HF Daily Papers' note
The authors introduce robot-centric pointmaps: image-like inputs whose pixels store 3D scene coordinates in the robot frame. The format keeps the dense grid expected by pretrained 2D VLA models, so it can be added with minimal architecture changes. On RoboCasa, the method improves pi0.5 and SmolVLA and beats camera-viewpoint and 3D-aware baselines. In real-robot tests, the gain over RGB-only control is larger when the camera is moved to an unseen placement. HF Daily Papers' note
score 5