Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
Ego2Act tests whether video models can actually carry out a high-level manipulation goal, not just render plausible motion.
The benchmark contains 2,640 egocentric videos covering 110 real-world daily tasks with clutter and multi-step complexity. Models are given an initial scene image and a goal, then judged on whether the generated hand-object video completes the task realistically. The paper also introduces Ego2ActJudge, a reference-free evaluator that better matches human consensus on task completion and physics plausibility than related baselines. The authors report that current models often skip steps, lose required intermediate states, and fail on fine physical dynamics in complex manipulation. HF Daily Papers' note
The benchmark contains 2,640 egocentric videos covering 110 real-world daily tasks with clutter and multi-step complexity. Models are given an initial scene image and a goal, then judged on whether the generated hand-object video completes the task realistically. The paper also introduces Ego2ActJudge, a reference-free evaluator that better matches human consensus on task completion and physics plausibility than related baselines. The authors report that current models often skip steps, lose required intermediate states, and fail on fine physical dynamics in complex manipulation. HF Daily Papers' note
score 5