Grounded Action Model: 3D Grounding as a Foundation for Robotics
GAM makes 3D object grounding the control interface for robot manipulation, rather than something learned indirectly from demos.
The paper says prompts in language, points, or boxes are converted into a shared object-centric representation with visual features and metric geometry. That representation is mixed with robot state history to predict action chunks, and can also sit under a higher-level planner for longer tasks. Reported results include 55.3% average success across 50 RoboTwin 2.0 tasks and 61% on LIBERO-PRO across 16 perturbation settings. On real robots, the model held up better under visual shift and completed long-horizon steps when paired with a Molmo2 planner. HF Daily Papers' note
The paper says prompts in language, points, or boxes are converted into a shared object-centric representation with visual features and metric geometry. That representation is mixed with robot state history to predict action chunks, and can also sit under a higher-level planner for longer tasks. Reported results include 55.3% average success across 50 RoboTwin 2.0 tasks and 61% on LIBERO-PRO across 16 perturbation settings. On real robots, the model held up better under visual shift and completed long-horizon steps when paired with a Molmo2 planner. HF Daily Papers' note
score 5