GRACE: Generation-aware latent compression for efficient video generation
GRACE claims an 8x token cut and 11.1x lower latency for Wan2.1-I2V-14B while preserving VBench quality.
The paper targets the bottleneck created when video diffusion models process large latent token streams. Its method compresses a pretrained video autoencoder without breaking compatibility with the pretrained DiT, using a frozen base latent plus a learned residual latent. The authors also align the compressed latent inside the frozen DiT’s feature space, then apply lightweight DiT fine-tuning with asymmetric denoising. HF Daily Papers' note
The paper targets the bottleneck created when video diffusion models process large latent token streams. Its method compresses a pretrained video autoencoder without breaking compatibility with the pretrained DiT, using a frozen base latent plus a learned residual latent. The authors also align the compressed latent inside the frozen DiT’s feature space, then apply lightweight DiT fine-tuning with asymmetric denoising. HF Daily Papers' note
score 4