UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
A pretrained semantic ViT can keep fine visual detail if its patch input is reparameterized.
The paper argues that the loss of pixel-level detail comes from the original patch parameterization, not from frozen Transformer blocks themselves. Its Patch Reparameterization keeps the semantic path while adding reconstruction-aware patch embeddings into the same ViT. The authors scale this into UniSpace, an 8B Mixture-of-Transformer-Experts model for understanding, generation, and editing without a separate VAE pathway. They report system evaluations showing practical text-to-image generation and instruction-based image editing. HF Daily Papers' note
The paper argues that the loss of pixel-level detail comes from the original patch parameterization, not from frozen Transformer blocks themselves. Its Patch Reparameterization keeps the semantic path while adding reconstruction-aware patch embeddings into the same ViT. The authors scale this into UniSpace, an 8B Mixture-of-Transformer-Experts model for understanding, generation, and editing without a separate VAE pathway. They report system evaluations showing practical text-to-image generation and instruction-based image editing. HF Daily Papers' note
score 5