UniWorld-Design: From Pixel Generation to Layer-Native Design
UniWorld-Design turns image generation into editable semantic RGBA layers, not a finished flat picture.
The paper describes two models: T2RGBA for generating standalone transparent assets from text, and I2L for decomposing a finished image into ordered semantic layers. I2L can follow global and per-layer prompts for decomposition, recursive decomposition, or targeted extraction. The authors say its layers model complete objects, so they remain useful when moved or removed. On Crello, I2L reports a 37% lower per-layer RGB L1 error and a 34% relative Alpha Soft IoU gain over Qwen-Image-Layered; T2RGBA reports the top CLIP Score against LayerDiffuse and OmniAlpha. HF Daily Papers' note
The paper describes two models: T2RGBA for generating standalone transparent assets from text, and I2L for decomposing a finished image into ordered semantic layers. I2L can follow global and per-layer prompts for decomposition, recursive decomposition, or targeted extraction. The authors say its layers model complete objects, so they remain useful when moved or removed. On Crello, I2L reports a 37% lower per-layer RGB L1 error and a 34% relative Alpha Soft IoU gain over Qwen-Image-Layered; T2RGBA reports the top CLIP Score against LayerDiffuse and OmniAlpha. HF Daily Papers' note
score 5