DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence
DC-SAE is pitched as a high-compression tokenizer that keeps diffusion training from slowing down.
The paper combines semantic encoders for compact latents with a pixel-level encoder meant to preserve reconstruction detail. On 512x512 ImageNet, it reports 32x spatial compression, 29.79 PSNR, and 3.37 gFID, beating DC-AE on those metrics while keeping comparable throughput. The authors also report faster diffusion convergence, plus text-to-image results from a 1.6B-parameter DiT at 1024x1024: 0.84 on GenEval and 86.007 on DPG-Bench. Source: HF Daily Papers' note.
The paper combines semantic encoders for compact latents with a pixel-level encoder meant to preserve reconstruction detail. On 512x512 ImageNet, it reports 32x spatial compression, 29.79 PSNR, and 3.37 gFID, beating DC-AE on those metrics while keeping comparable throughput. The authors also report faster diffusion convergence, plus text-to-image results from a 1.6B-parameter DiT at 1024x1024: 0.84 on GenEval and 86.007 on DPG-Bench. Source: HF Daily Papers' note.
score 5