Megadose AI progress, ranked and analyzed.

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

· HF Daily Papers ·
The paper argues that VLA models generalize better when scene geometry is expressed in the robot’s own coordinate frame.

The authors introduce robot-centric pointmaps: image-like inputs whose pixels store 3D scene coordinates in the robot frame. The format keeps the dense grid expected by pretrained 2D VLA models, so it can be added with minimal architecture changes. On RoboCasa, the method improves pi0.5 and SmolVLA and beats camera-viewpoint and 3D-aware baselines. In real-robot tests, the gain over RGB-only control is larger when the camera is moved to an unseen placement. HF Daily Papers' note

score 5

Categories: Research