Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
Sci-VBench tests whether video models can get scientific cause and sequence right, not just look realistic.
The benchmark has 1,253 expert-annotated prompts across 60 subjects in science, healthcare, social sciences, and engineering. The authors evaluate 16 proprietary and open-source models with a rubric for prompt grounding and scientific or causal correctness. They report that visual-quality scores are close across systems, while scientific reasoning performance varies widely, especially between proprietary and open-source models. ArXiv · AI/CL/LG's note
The benchmark has 1,253 expert-annotated prompts across 60 subjects in science, healthcare, social sciences, and engineering. The authors evaluate 16 proprietary and open-source models with a rubric for prompt grounding and scientific or causal correctness. They report that visual-quality scores are close across systems, while scientific reasoning performance varies widely, especially between proprietary and open-source models. ArXiv · AI/CL/LG's note
score 5