Megadose AI progress, ranked and analyzed.

Scaling Native Multimodal Pre-Training From Scratch

· HF Daily Papers ·
The paper claims native multimodal models follow predictable scaling laws, but the optimal mix changes sharply with data composition.

The authors train transformer vision-language models from scratch under fixed compute budgets and find objective loss tracks a compute law. Language learning stays relatively stable across different multimodal data ratios. Multimodal learning does not: text-heavy mixtures only become compute-efficient at larger model sizes. The paper also reports positive cross-modal transfer, including gains in pure-text spatial reasoning and multimodal in-context learning. HF Daily Papers' note

score 5

Categories: Research