Editable Visual Design
The paper proposes a coding-agent workflow that generates editable HTML/CSS designs instead of flattened image outputs.
The system uses a vision-language model for planning and aesthetic judgment, and calls an image model to synthesize isolated visual assets when needed. It then writes native HTML/CSS and refines the result against rendered visual feedback. The authors say the output keeps layers decoupled and text real, so users can drag and adjust elements in a GUI. They report validations on posters, infographics, and other design scenarios. ArXiv · AI/CL/LG's note
The system uses a vision-language model for planning and aesthetic judgment, and calls an image model to synthesize isolated visual assets when needed. It then writes native HTML/CSS and refines the result against rendered visual feedback. The authors say the output keeps layers decoupled and text real, so users can drag and adjust elements in a GUI. They report validations on posters, infographics, and other design scenarios. ArXiv · AI/CL/LG's note
score 5