WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
WorldExam tests whether video “world models” react plausibly to scene conditions, not just whether they look good or follow prompts.
The benchmark has 1,474 cases across eight tasks and four levels: visual quality, control adherence, spatial consistency, and world reactivity. Its hardest layer asks models to generate scene-conditioned reactions and goal-directed behavior beyond explicit input instructions. In tests of 20 models, camera-driven systems handled camera control well but lacked dynamic interaction, action-driven systems controlled subjects but often left scenes unresponsive, and language-driven systems interacted better while struggling with complex controls. The authors report that no model combined broad coverage with consistently strong performance. HF Daily Papers' note
The benchmark has 1,474 cases across eight tasks and four levels: visual quality, control adherence, spatial consistency, and world reactivity. Its hardest layer asks models to generate scene-conditioned reactions and goal-directed behavior beyond explicit input instructions. In tests of 20 models, camera-driven systems handled camera control well but lacked dynamic interaction, action-driven systems controlled subjects but often left scenes unresponsive, and language-driven systems interacted better while struggling with complex controls. The authors report that no model combined broad coverage with consistently strong performance. HF Daily Papers' note
score 5