VGI-BENCH: Probing Visual Intelligence in Video Generation Models
The benchmark’s top tested video model reached only 51.0% on visually grounded reasoning tasks.
VGI-Bench defines 27 tasks and 810 instances to test whether video generators can produce valid evolving processes, not just plausible end frames. The authors say current systems can solve some tasks but are still far from reliable. Their analysis points to failure modes, sensitivity to input conditions, limited transfer from synthetic fine-tuning, and denoising steps that tend to refine early guesses rather than fix reasoning mistakes. Source: HF Daily Papers' note
VGI-Bench defines 27 tasks and 810 instances to test whether video generators can produce valid evolving processes, not just plausible end frames. The authors say current systems can solve some tasks but are still far from reliable. Their analysis points to failure modes, sensitivity to input conditions, limited transfer from synthetic fine-tuning, and denoising steps that tend to refine early guesses rather than fix reasoning mistakes. Source: HF Daily Papers' note
score 4