SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
The paper tests whether video models can produce the requested end state, not just plausible motion.
SemComp-Bench frames video generation as semantic task completion: the output must achieve an intended outcome while staying grounded in task-relevant semantics from a reference image. The authors build SemComp-Data across six domains, pairing reference images with detailed and brief instructions plus outcome-focused clips. Their protocol uses a vision-language model to answer binary questions and reports Outcome Achievement and Generation Reliability scores. Experiments on representative models find that this outcome-plus-grounding requirement remains difficult. HF Daily Papers' note
SemComp-Bench frames video generation as semantic task completion: the output must achieve an intended outcome while staying grounded in task-relevant semantics from a reference image. The authors build SemComp-Data across six domains, pairing reference images with detailed and brief instructions plus outcome-focused clips. Their protocol uses a vision-language model to answer binary questions and reports Outcome Achievement and Generation Reliability scores. Experiments on representative models find that this outcome-plus-grounding requirement remains difficult. HF Daily Papers' note
score 5