Abra: Scaling Diffusion Image Training
Abra finds diffusion training follows scaling laws, but wants far more data per parameter than LLM rules suggest.
The paper studies a controlled family of flow-matching text-to-image transformers trained from $10^{19}$ to $10^{22}$ FLOPs. Its main result is that compute-optimal diffusion training lands around 200 image tokens per parameter, about 10x the Chinchilla prescription for language models. The authors also report that diffusion models tolerate overtraining better than language models, making extra data the safer bet versus simply scaling model size. The same predictability is said to extend beyond loss into generation metrics, CFG settings, representation quality, and training-curve shape. HF Daily Papers' note
The paper studies a controlled family of flow-matching text-to-image transformers trained from $10^{19}$ to $10^{22}$ FLOPs. Its main result is that compute-optimal diffusion training lands around 200 image tokens per parameter, about 10x the Chinchilla prescription for language models. The authors also report that diffusion models tolerate overtraining better than language models, making extra data the safer bet versus simply scaling model size. The same predictability is said to extend beyond loss into generation metrics, CFG settings, representation quality, and training-curve shape. HF Daily Papers' note
score 7