Megadose Built for builders and researchers.

Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

· HF Daily Papers ·
The paper’s central result is negative: pixel-space diffusion did not improve the tested downstream tasks.

Iris-3B is a 3B-parameter pixel-space text-to-image transformer trained through a 256-to-1024 curriculum. The authors also converted FLUX.2 Klein base 4B from latent space to pixel space, then fine-tuned both model families for monocular depth and restoration/super-resolution. In their reported tests, Iris-3B matched the latent FLUX.2 Klein on depth, while the converted pixel model trailed; on 4x DIV2K restoration, neither pixel model beat the latent fine-tune. The paper still reports that Iris-3B scales pixel-space pretraining to 3B parameters and reaches text-to-image quality competitive with latent models at 1024². HF Daily Papers' note

score 4

Categories: Research