UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
UltraText Bench tests whether image models can keep long bilingual text accurate, legible, and correctly placed across crowded scenes.
The benchmark includes 432 human-reviewed prompts across 24 scene categories, split evenly between English and Chinese. Each prompt specifies exact strings for four to twelve text regions, plus references for content, placement, and visual attributes. The authors score outputs with Q-Judger on fidelity, clarity, spatial quality, and scene quality, then compare 24 model configurations. Their reported results show tradeoffs: Z-Image-Turbo improves clarity over Z-Image-Base but loses fidelity, while Qwen-Image-2512 drops sharply as English prompt difficulty rises. HF Daily Papers' note
The benchmark includes 432 human-reviewed prompts across 24 scene categories, split evenly between English and Chinese. Each prompt specifies exact strings for four to twelve text regions, plus references for content, placement, and visual attributes. The authors score outputs with Q-Judger on fidelity, clarity, spatial quality, and scene quality, then compare 24 model configurations. Their reported results show tradeoffs: Z-Image-Turbo improves clarity over Z-Image-Base but loses fidelity, while Qwen-Image-2512 drops sharply as English prompt difficulty rises. HF Daily Papers' note
score 4