Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
The paper tests image generators by making them answer spatial tasks in pixels, then converting those marks back into benchmark metrics.
ProVisE constrains models to point, mark, or draw instead of forcing spatial answers into coordinates or text. The authors pair it with SpatialGen-Bench, a 470-sample diagnostic set covering 14 spatial subtasks. In their unified evaluation, image-generation models are competitive when pixel-space answers fit the task, while text-output VLMs still lead on compositional spatial reasoning. HF Daily Papers' note
ProVisE constrains models to point, mark, or draw instead of forcing spatial answers into coordinates or text. The authors pair it with SpatialGen-Bench, a 470-sample diagnostic set covering 14 spatial subtasks. In their unified evaluation, image-generation models are competitive when pixel-space answers fit the task, while text-output VLMs still lead on compositional spatial reasoning. HF Daily Papers' note
score 4