SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments
SeededGrasp splits language understanding from grasp execution by having the VLM choose a 3D seed point first.
The paper says that seed then conditions a lightweight grasp-generation model, avoiding expensive end-to-end VLM-and-grasp training. The authors position this as a way to handle cluttered tabletop scenes across multiple robot embodiments. They also release a multi-embodiment dataset with more than 2.5M grasps. Reported results are 72% success in simulation and 78% in real-world grasping experiments. HF Daily Papers' note
The paper says that seed then conditions a lightweight grasp-generation model, avoiding expensive end-to-end VLM-and-grasp training. The authors position this as a way to handle cluttered tabletop scenes across multiple robot embodiments. They also release a multi-embodiment dataset with more than 2.5M grasps. Reported results are 72% success in simulation and 78% in real-world grasping experiments. HF Daily Papers' note
score 4