Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
The paper argues that image tokenizers should be judged by how their tokens behave alongside text, not just by reconstruction or single-task scores.
The authors build a controlled pure-autoregressive testbed for multimodal continual pretraining across text, image, text-to-image, and image-to-text prediction. They find task-specific losses scale differently and can rank tokenizers differently. Image-to-text loss, measured over a shared text vocabulary, is presented as a more consistent signal across tokenizers than text-to-image loss. The paper also reports that stronger reconstruction does not necessarily mean better downstream performance, and that tokenizer choice can affect text modeling under joint optimization. ArXiv · AI/CL/LG's note
The authors build a controlled pure-autoregressive testbed for multimodal continual pretraining across text, image, text-to-image, and image-to-text prediction. They find task-specific losses scale differently and can rank tokenizers differently. Image-to-text loss, measured over a shared text vocabulary, is presented as a more consistent signal across tokenizers than text-to-image loss. The paper also reports that stronger reconstruction does not necessarily mean better downstream performance, and that tokenizer choice can affect text modeling under joint optimization. ArXiv · AI/CL/LG's note
score 5