FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
FilmBench tests video models against film-school craft, and the scores drop when prompts demand real multi-shot cinema.
The benchmark builds 1,169 prompts from award-winning film clips across 20 genres, with professional directors selecting the references. Most prompts are multi-shot, and the evaluation uses a Cinematic Language taxonomy spanning axes, components, and sub-metrics rather than generic quality checks. Its FilmOps evaluator closely matches human model rankings, with Spearman scores of 0.95 for text-to-video and 0.96 for reference-to-video. The paper says leading models show persistent weaknesses in dynamic aesthetics and lose more ground on multi-shot tasks, especially the weaker systems. HF Daily Papers' note
The benchmark builds 1,169 prompts from award-winning film clips across 20 genres, with professional directors selecting the references. Most prompts are multi-shot, and the evaluation uses a Cinematic Language taxonomy spanning axes, components, and sub-metrics rather than generic quality checks. Its FilmOps evaluator closely matches human model rankings, with Spearman scores of 0.95 for text-to-video and 0.96 for reference-to-video. The paper says leading models show persistent weaknesses in dynamic aesthetics and lose more ground on multi-shot tasks, especially the weaker systems. HF Daily Papers' note
score 5