H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
H2R-Bench tests whether video world models can turn human manipulation footage into robot-specific demonstrations.
The benchmark pairs egocentric human demonstration videos with target robot embodiment constraints and grounded annotations for goals, actions, contacts, and object responses. It scores generated videos on task completion, action events, functional contact transfer, embodiment correctness, and video quality. The authors evaluate eleven video generation models across six manipulation families and two robot embodiments. Their finding is that current models still often miss embodiment consistency, functional interaction, and execution. HF Daily Papers' note
The benchmark pairs egocentric human demonstration videos with target robot embodiment constraints and grounded annotations for goals, actions, contacts, and object responses. It scores generated videos on task completion, action events, functional contact transfer, embodiment correctness, and video quality. The authors evaluate eleven video generation models across six manipulation families and two robot embodiments. Their finding is that current models still often miss embodiment consistency, functional interaction, and execution. HF Daily Papers' note
score 4