VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
VTR-Bench tests whether video models can put readable, correct text into generated scenes.
The benchmark uses 300 prompts across five scenario categories, including advertisements and scientific videos. Its evaluation separates text fidelity from scene and motion requirements, using automated checks aligned with human judgment. The authors report that 11 state-of-the-art models still struggle, with the best overall word error rate at 0.250. They also propose a keyframe-guided agentic framework to refine generations through visual feedback. HF Daily Papers' note
The benchmark uses 300 prompts across five scenario categories, including advertisements and scientific videos. Its evaluation separates text fidelity from scene and motion requirements, using automated checks aligned with human judgment. The authors report that 11 state-of-the-art models still struggle, with the best overall word error rate at 0.250. They also propose a keyframe-guided agentic framework to refine generations through visual feedback. HF Daily Papers' note
score 4