Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
The paper argues that multimodal models work best when vision and language are trained together from the start.
The authors report controlled experiments on synthetic and large-scale datasets to trace how language, visual understanding, and visual generation transfer knowledge across modalities. They say data complexity helps determine whether modalities reinforce each other or compete. The study favors shared attention and normalization with modality-specific feed-forward layers, and warns that delayed visual integration can make models lean on language priors. Its proposed recipes reach strong generative performance with 5% of the compute budget, then are validated on 13.5B MoE models trained on 2T tokens. HF Daily Papers' note
The authors report controlled experiments on synthetic and large-scale datasets to trace how language, visual understanding, and visual generation transfer knowledge across modalities. They say data complexity helps determine whether modalities reinforce each other or compete. The study favors shared attention and normalization with modality-specific feed-forward layers, and warns that delayed visual integration can make models lean on language priors. Its proposed recipes reach strong generative performance with 5% of the compute budget, then are validated on 13.5B MoE models trained on 2T tokens. HF Daily Papers' note
score 5