Megadose AI progress, ranked and analyzed.

Looped Diffusion Transformer

· ArXiv · AI/CL/LG ·
Looped-DiT scales text-to-image generation by rerunning shared Transformer blocks instead of adding parameters or denoising steps.

The paper says naive looping does not reliably improve quality because intermediate loops are weakly supervised and attention updates can erode local detail. Its fix combines deep supervision across loops with self-modulating attention. In matched settings, the authors report that a 260M-parameter looped model beats a 6.5x larger model while using 4.9x less inference compute. They also find deeper loops can correct earlier mistakes under a fixed inference budget. ArXiv · AI/CL/LG's note

score 4

Categories: Research