Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
The paper pins some physical failures in generated video on attention heads that lock motion too early.
The authors trace “motion planning” to early denoising stages in text-to-video diffusion models. They identify a subset of attention heads that steer candidate trajectories, then argue RoPE makes self-attention decay too sharply across space. That premature decay can freeze objects into implausible positions and block better motion in nearby frames. Their proposed fix scales RoPE frequency across denoising steps, with experiments reported as improving physical commonsense in generated videos. HF Daily Papers' note
The authors trace “motion planning” to early denoising stages in text-to-video diffusion models. They identify a subset of attention heads that steer candidate trajectories, then argue RoPE makes self-attention decay too sharply across space. That premature decay can freeze objects into implausible positions and block better motion in nearby frames. Their proposed fix scales RoPE frequency across denoising steps, with experiments reported as improving physical commonsense in generated videos. HF Daily Papers' note
score 5