Megadose Built for builders and researchers.

V-RAE: Rethinking Video Latent Spaces for Generation

· HF Daily Papers ·
V-RAE argues that video latents built from frozen semantic encoders can serve generation better than reconstruction-first VAEs.

The paper adds a lightweight temporal pooling module to compress frozen vision-model features while keeping semantic structure. Its best reported results include 2.13 rFVD on K600 and gFVD scores of 117.86 on UCF101 and 19.16 on K600 under matched generation settings. The authors say the model converges up to 6x faster and retains more semantic information than conventional video tokenizer latents. They also introduce tFVD as a temporal-coherence diagnostic for judging generative utility. Source: HF Daily Papers' note.

score 5

Categories: Research