Megadose AI progress, ranked and analyzed.

ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation

· HF Daily Papers ·
ViRDM removes the teacher-and-critic training stack from few-step causal video post-training.

The paper presents a generator-only recipe for representation distribution matching in low-latency autoregressive video diffusion. It targets three blockers the authors found in applying RDM to video: memory cost, video-specific optimization behavior, and weak temporal constraints. The reported setup uses truncated clean-exit supervision, a lightweight VAE decoder, staged vector-Jacobian products, and dynamics regularization. With 20 generator updates, it reaches 84.87 on VBench, 0.36 above the prior few-step causal baseline, using 16 A100 GPU-hours. HF Daily Papers' note

score 5

Categories: Research