Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
Boogu-Image-0.1 claims near closed-system image generation results with a roughly $400K base-model training bill.
The paper presents an open-source model family for text-to-image generation, fast inference, instruction editing, and Chinese-English text rendering. Its core claim is that stronger multimodal understanding, prompt rewriting, better data, and inference-time scaling can lift generation quality under tight compute limits. The authors say it matches or beats other open-source models on standard benchmarks and approaches leading closed-source systems. They release weights, code, and recipes under Apache 2.0. HF Daily Papers' note
The paper presents an open-source model family for text-to-image generation, fast inference, instruction editing, and Chinese-English text rendering. Its core claim is that stronger multimodal understanding, prompt rewriting, better data, and inference-time scaling can lift generation quality under tight compute limits. The authors say it matches or beats other open-source models on standard benchmarks and approaches leading closed-source systems. They release weights, code, and recipes under Apache 2.0. HF Daily Papers' note
score 6