SenseNova-U1.5: Towards Native Unified Visual Intelligence
SenseNova-U1.5 is an 8B unified multimodal model built to understand, reason over, edit, and generate images without separate encoder or VAE components.
The paper says the model uses spatially coherent patch reconstruction and trains up to native 4K resolution. Its post-training adds specialized experts for aesthetics, bilingual text rendering, infographics, and image editing, then distills them into one system. The authors report gains in image fidelity, text rendering, complex composition, multi-reference editing, and preserving identity, geometry, and untouched regions. They also say they will open-source training code for supervised fine-tuning, reinforcement learning, and on-policy distillation. HF Daily Papers' note
The paper says the model uses spatially coherent patch reconstruction and trains up to native 4K resolution. Its post-training adds specialized experts for aesthetics, bilingual text rendering, infographics, and image editing, then distills them into one system. The authors report gains in image fidelity, text rendering, complex composition, multi-reference editing, and preserving identity, geometry, and untouched regions. They also say they will open-source training code for supervised fine-tuning, reinforcement learning, and on-policy distillation. HF Daily Papers' note
score 6