Beyond the Current Scene: Event-Referential Grasping with Active View Selection
BeyondSCe lets a robot grasp objects referred to by a past interaction, even when the object is no longer visible.
The paper frames the task as event-referential grasping: a person asks for the object or part tied to an earlier event, not just a named visible item. BeyondSCe uses event history and the current scene to identify and localize the target, then chooses new camera views when occlusion blocks it. It runs zero-shot with pretrained models and a single wrist-mounted RGB-D camera. In real-robot tests, it reports 76% grasp success for visible targets and 77% for occluded ones, ahead of the strongest baselines cited in the abstract. HF Daily Papers' note
The paper frames the task as event-referential grasping: a person asks for the object or part tied to an earlier event, not just a named visible item. BeyondSCe uses event history and the current scene to identify and localize the target, then chooses new camera views when occlusion blocks it. It runs zero-shot with pretrained models and a single wrist-mounted RGB-D camera. In real-robot tests, it reports 76% grasp success for visible targets and 77% for occluded ones, ahead of the strongest baselines cited in the abstract. HF Daily Papers' note
score 5