Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
Sci-VBench tests whether video models can show scientific cause and effect, not just make plausible-looking clips.
The benchmark has 1,253 expert-annotated examples across 60 subjects in science, healthcare, social sciences, and engineering. Its rubric checks prompt grounding plus scientific and causal correctness. The authors say human non-experts and MLLM judges can agree fairly well with expert ratings under that protocol. Across 16 models, visual-quality scores were close, but scientific reasoning performance varied sharply, with proprietary systems ahead of open-source ones. HF Daily Papers' note
The benchmark has 1,253 expert-annotated examples across 60 subjects in science, healthcare, social sciences, and engineering. Its rubric checks prompt grounding plus scientific and causal correctness. The authors say human non-experts and MLLM judges can agree fairly well with expert ratings under that protocol. Across 16 models, visual-quality scores were close, but scientific reasoning performance varied sharply, with proprietary systems ahead of open-source ones. HF Daily Papers' note
score 5