Megadose AI progress, ranked and analyzed.

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

· ArXiv · AI/CL/LG ·
Sci-VBench tests whether video models can get scientific cause and sequence right, not just look realistic.

The benchmark has 1,253 expert-annotated prompts across 60 subjects in science, healthcare, social sciences, and engineering. The authors evaluate 16 proprietary and open-source models with a rubric for prompt grounding and scientific or causal correctness. They report that visual-quality scores are close across systems, while scientific reasoning performance varies widely, especially between proprietary and open-source models. ArXiv · AI/CL/LG's note

score 5

Categories: Research