Megadose Built for builders and researchers.

HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution

· ArXiv · AI/CL/LG ·
The benchmark finds humanoids can pick the right tool more often than they can finish the job with it.

HumanoidToolBench tests 18 tasks across tool selection, manipulation, and mobile execution. The authors also release ToolBook, with 3.1k demonstrations from simulation and a real Unitree G1. Their evaluations show a clear drop from choosing a suitable tool to actually completing the task. GR00T N1.7 probes further show weaker accuracy on unseen tools and cases where execution continues despite unrelated instructions. ArXiv · AI/CL/LG's note

score 5

Categories: Research