Megadose Built for builders and researchers.

Abra: Scaling Diffusion Image Training

· HF Daily Papers ·
Abra finds diffusion training follows scaling laws, but wants far more data per parameter than LLM rules suggest.

The paper studies a controlled family of flow-matching text-to-image transformers trained from $10^{19}$ to $10^{22}$ FLOPs. Its main result is that compute-optimal diffusion training lands around 200 image tokens per parameter, about 10x the Chinchilla prescription for language models. The authors also report that diffusion models tolerate overtraining better than language models, making extra data the safer bet versus simply scaling model size. The same predictability is said to extend beyond loss into generation metrics, CFG settings, representation quality, and training-curve shape. HF Daily Papers' note

score 7

Categories: Research