HARMONY: Hierarchical Agentic Reasoning for MONocular Image-to-Scene Synthesis
HARMONY turns one indoor image into a structured 3D scene by combining VLM placement reasoning with point-cloud refinement.
The paper frames single-image scene reconstruction as a failure point for either semantic agents or geometry models alone. HARMONY starts from an empty floorplan, calibrates the camera, infers room layout, then places objects in order: wall-mounted items, furniture, and dependent decorations. After each VLM placement, geometry estimates refine alignment with the input image. The authors report stronger results than evaluated reconstruction baselines, with qualitative comparisons claiming more faithful arrangements than GPT-6 Astra. HF Daily Papers' note
The paper frames single-image scene reconstruction as a failure point for either semantic agents or geometry models alone. HARMONY starts from an empty floorplan, calibrates the camera, infers room layout, then places objects in order: wall-mounted items, furniture, and dependent decorations. After each VLM placement, geometry estimates refine alignment with the input image. The authors report stronger results than evaluated reconstruction baselines, with qualitative comparisons claiming more faithful arrangements than GPT-6 Astra. HF Daily Papers' note
score 4