Mitigating Compounding Error via Video Representation Regularization
The paper ties long-video generation drift to hidden representation collapse, then tests a regularizer meant to keep those representations stable.
Chen, Zhang, and Wang say sliding-window autoregressive video diffusion models lose frame quality as errors accumulate over time. Their analysis finds the effective rank of hidden representations drops sharply when generation drift begins. They report that simply scaling training data did not improve resistance to drift. Their proposed video representation regularization improves VBench Aesthetic Quality from 38.65 to 55.56 and Imaging Quality from 44.37 to 72.08 versus Diffusion Forcing. ArXiv · AI/CL/LG's note
Chen, Zhang, and Wang say sliding-window autoregressive video diffusion models lose frame quality as errors accumulate over time. Their analysis finds the effective rank of hidden representations drops sharply when generation drift begins. They report that simply scaling training data did not improve resistance to drift. Their proposed video representation regularization improves VBench Aesthetic Quality from 38.65 to 55.56 and Imaging Quality from 44.37 to 72.08 versus Diffusion Forcing. ArXiv · AI/CL/LG's note
score 5