Structured Residual Connectivity Matters for Diffusion Transformers
The paper argues DiTs train and synthesize better when residual paths actively retrieve earlier features instead of treating depth as one pooled stream.
The authors report that DiT representations show a preference for early-layer reuse and symmetric layer guidance. Their proposed connectivity adds local residual links plus long-range cross-depth paths, letting blocks attend to earlier spatial and semantic cues. In experiments, the approach cut training iterations by up to 1.73x and added under 0.1% parameters. It improved a REPA-XL/2 model from 5.9 to 4.34 FID without guidance, and reached 1.39 FID with classifier-free guidance. HF Daily Papers' note
The authors report that DiT representations show a preference for early-layer reuse and symmetric layer guidance. Their proposed connectivity adds local residual links plus long-range cross-depth paths, letting blocks attend to earlier spatial and semantic cues. In experiments, the approach cut training iterations by up to 1.73x and added under 0.1% parameters. It improved a REPA-XL/2 model from 5.9 to 4.34 FID without guidance, and reached 1.39 FID with classifier-free guidance. HF Daily Papers' note
score 4