Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
LDR is presented as a video world model that reasons over latent motion dynamics instead of only fitting plausible pixels.
The paper says Latent Dynamics Reasoning treats frame-to-frame change as kinematic integration, with the model learning higher-order residuals for rollout. It is tested on five controlled physics tasks, including collisions, bouncing, and parabolic motion, with emphasis on out-of-distribution cases. The authors report more than a 20x smaller in/out-of-distribution error gap than a video diffusion baseline, while using 26x fewer parameters and running 143x faster. They also claim it generalizes under severe shifts, such as predicting a blue square moving right-to-left after training on red balls moving left-to-right. HF Daily Papers' note
The paper says Latent Dynamics Reasoning treats frame-to-frame change as kinematic integration, with the model learning higher-order residuals for rollout. It is tested on five controlled physics tasks, including collisions, bouncing, and parabolic motion, with emphasis on out-of-distribution cases. The authors report more than a 20x smaller in/out-of-distribution error gap than a video diffusion baseline, while using 26x fewer parameters and running 143x faster. They also claim it generalizes under severe shifts, such as predicting a blue square moving right-to-left after training on red balls moving left-to-right. HF Daily Papers' note
score 5