Megadose AI progress, ranked and analyzed.

HumanCLAW: Can Vision-Language Models Act Through a Body?

· HF Daily Papers ·
The benchmark isolates VLM decision-making from motor failures, and the best tested model still solved only 16.8% of episodes.

HumanCLAW has a harnessed VLM choose atomic body commands while a controller handles short chunks of full-body motion under gravity and collisions. Its benchmark covers 1,218 long-horizon egocentric find-navigate-interact episodes in 41 indoor scenes. The paper says target recognition is not the main failure point; models instead lose track of their own body, position, progress, and collisions. HF Daily Papers' note

score 4

Categories: Research