Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
The paper argues that image tokenizers should be judged by how their tokens train alongside text, not just by reconstruction or single-task scores.
The authors build a controlled pure-autoregressive setup and track validation losses across text, image, text-to-image, and image-to-text prediction. They find those losses scale differently and can rank tokenizers differently by task. Image-to-text loss, measured over a shared text vocabulary, is presented as the more consistent signal across tokenizers and is linked to later generation and visual-understanding results after finetuning. Better reconstruction alone does not guarantee stronger multimodal performance, and tokenizer choice can even affect text modeling under joint training. Source: HF Daily Papers' note.
The authors build a controlled pure-autoregressive setup and track validation losses across text, image, text-to-image, and image-to-text prediction. They find those losses scale differently and can rank tokenizers differently by task. Image-to-text loss, measured over a shared text vocabulary, is presented as the more consistent signal across tokenizers and is linked to later generation and visual-understanding results after finetuning. Better reconstruction alone does not guarantee stronger multimodal performance, and tokenizer choice can even affect text modeling under joint training. Source: HF Daily Papers' note.
score 5