HumanCLAW: Can Vision-Language Models Act Through a Body?
The benchmark isolates VLM decision-making from motor failures, and the best tested model still solved only 16.8% of episodes.
HumanCLAW has a harnessed VLM choose atomic body commands while a controller handles short chunks of full-body motion under gravity and collisions. Its benchmark covers 1,218 long-horizon egocentric find-navigate-interact episodes in 41 indoor scenes. The paper says target recognition is not the main failure point; models instead lose track of their own body, position, progress, and collisions. HF Daily Papers' note
HumanCLAW has a harnessed VLM choose atomic body commands while a controller handles short chunks of full-body motion under gravity and collisions. Its benchmark covers 1,218 long-horizon egocentric find-navigate-interact episodes in 41 indoor scenes. The paper says target recognition is not the main failure point; models instead lose track of their own body, position, progress, and collisions. HF Daily Papers' note
score 4