MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment
The paper claims object-level pretraining can beat far larger image-text alignment runs.
MMCS replaces textual entities with their matching visual objects during pretraining, forcing local grounding instead of relying on a whole-image representation. The authors say this reduces referential ambiguity when an image contains multiple objects and entities. They report a 773K-sample synthetic dataset, with 50K MMCS samples matching or surpassing models trained on 600K image-text pairs. ArXiv · AI/CL/LG's note
MMCS replaces textual entities with their matching visual objects during pretraining, forcing local grounding instead of relying on a whole-image representation. The authors say this reduces referential ambiguity when an image contains multiple objects and entities. They report a 773K-sample synthetic dataset, with 50K MMCS samples matching or surpassing models trained on 600K image-text pairs. ArXiv · AI/CL/LG's note
score 5