Megadose AI progress, ranked and analyzed.

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

· ArXiv · AI/CL/LG ·
The paper argues that image tokenizers should be judged by how their tokens behave alongside text, not just by reconstruction or single-task scores.

The authors build a controlled pure-autoregressive testbed for multimodal continual pretraining across text, image, text-to-image, and image-to-text prediction. They find task-specific losses scale differently and can rank tokenizers differently. Image-to-text loss, measured over a shared text vocabulary, is presented as a more consistent signal across tokenizers than text-to-image loss. The paper also reports that stronger reconstruction does not necessarily mean better downstream performance, and that tokenizer choice can affect text modeling under joint optimization. ArXiv · AI/CL/LG's note

score 5

Categories: Research