TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows
TerraVis tests generated images for whether they make physical and spatial sense, not just whether they look good.
The paper defines 18 types of world-consistency failures across objects, interactions, and scenes. Its workflow uses an MLLM to decide whether an image can be evaluated, detect violations, grade them as minor or major, and produce a consistency score. On two common benchmarks, TerraVis correlated more strongly with human judgments of world consistency than existing metrics. The authors also report that models scoring well on conventional measures can still make substantial real-world plausibility errors. HF Daily Papers' note
The paper defines 18 types of world-consistency failures across objects, interactions, and scenes. Its workflow uses an MLLM to decide whether an image can be evaluated, detect violations, grade them as minor or major, and produce a consistency score. On two common benchmarks, TerraVis correlated more strongly with human judgments of world consistency than existing metrics. The authors also report that models scoring well on conventional measures can still make substantial real-world plausibility errors. HF Daily Papers' note
score 4