Multimodal Thinking with Renderable Programs
SVGLM uses SVG primitives as an intermediate reasoning medium so vision-language models can generate and inspect images while they reason.
The paper argues that raster or latent image representations make visual reasoning harder to trace. Its framework treats SVG as both text instructions and renderable image description, aiming for compact and interpretable image-in-the-loop reasoning. The authors also describe a curated SVG-based image editing dataset and a tuning setup for open-source VLMs. In experiments on a math reasoning benchmark, they report stronger SVG generation and “think-with-image” behavior.
ArXiv · AI/CL/LG's note
The paper argues that raster or latent image representations make visual reasoning harder to trace. Its framework treats SVG as both text instructions and renderable image description, aiming for compact and interpretable image-in-the-loop reasoning. The authors also describe a curated SVG-based image editing dataset and a tuning setup for open-source VLMs. In experiments on a math reasoning benchmark, they report stronger SVG generation and “think-with-image” behavior.
ArXiv · AI/CL/LG's note
score 5