Megadose AI progress, ranked and analyzed.

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

· HF Daily Papers ·
Sci-VBench tests whether video models can show scientific cause and effect, not just make plausible-looking clips.

The benchmark has 1,253 expert-annotated examples across 60 subjects in science, healthcare, social sciences, and engineering. Its rubric checks prompt grounding plus scientific and causal correctness. The authors say human non-experts and MLLM judges can agree fairly well with expert ratings under that protocol. Across 16 models, visual-quality scores were close, but scientific reasoning performance varied sharply, with proprietary systems ahead of open-source ones. HF Daily Papers' note

score 5

Categories: Research