Megadose AI progress, ranked and analyzed.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

· ArXiv · AI/CL/LG ·
Early joint training beat late alignment in the paper’s controlled multimodal pretraining tests.

The authors report four findings: cross-modal knowledge transfer is asymmetric, modality synergy depends heavily on data complexity, and architecture choices can reduce competition. They say shared attention and normalization with modality-specific feed-forward layers helped promote synergy. Delayed integration produced what they call “vision laziness,” where models leaned on language priors. The recipes were then validated by training several 13.5B MoE models on 2T tokens, with strong generation results claimed at 5% of the compute budget. ArXiv · AI/CL/LG's note

score 6

Categories: Research