Megadose Built for builders and researchers.

An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models

· HF Daily Papers ·
A latent-to-pixel training route made pixel-space diffusion competitive while cutting inference time by 3.18x to 4.75x.

The paper says direct large-scale pretraining in pixel space converges much more slowly than latent-space training. Its proposed recipe learns generative priors in latent space, then moves into pixel space during post-training. The authors test transition choices including initialization, data mix, prediction target, decoder architecture, and noise schedule. They report that the resulting pixel-space models can match or outperform latent-space counterparts. HF Daily Papers' note

score 5

Categories: Research