LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
The paper claims a new open-source benchmark high with released weights, code, and training recipes.
LLaDA-Image pairs a 6B diffusion transformer trained from scratch with a frozen vision-language module built on LLaDA2.0-Mini. The authors emphasize image-only pre-training and mid-training before heavier use of paired image-text data. They also describe a Turbo version distilled for 2-4 sampling steps. On Qwen-Image-Bench, they report overall scores of 53.53 in English and 53.38 in Chinese. HF Daily Papers' note
LLaDA-Image pairs a 6B diffusion transformer trained from scratch with a frozen vision-language module built on LLaDA2.0-Mini. The authors emphasize image-only pre-training and mid-training before heavier use of paired image-text data. They also describe a Turbo version distilled for 2-4 sampling steps. On Qwen-Image-Bench, they report overall scores of 53.53 in English and 53.38 in Chinese. HF Daily Papers' note
score 6