Reasoning with Image Generation
ReImaGin uses image generation itself as the visual reasoning step, not just as an output tool.
The paper argues that text-only chain-of-thought falls short when a task requires changing or constructing visual representations. Its method lets a multimodal model issue natural-language commands to an image generator for operations such as removing occlusion or making a floorplan from separate room views. Across six visual reasoning tasks, including spatial reasoning and collision prediction, it beat text-only and specialist vision-tool baselines, with reported gains up to 25%. HF Daily Papers' note
The paper argues that text-only chain-of-thought falls short when a task requires changing or constructing visual representations. Its method lets a multimodal model issue natural-language commands to an image generator for operations such as removing occlusion or making a floorplan from separate room views. Across six visual reasoning tasks, including spatial reasoning and collision prediction, it beat text-only and specialist vision-tool baselines, with reported gains up to 25%. HF Daily Papers' note
score 5