Self-Supervised Learning of Structured Dynamics from Videos
The paper tests whether frozen image-transformer features can be made to separate camera motion from object motion in video.
The authors introduce a Structured Dynamics Model that predicts future features while splitting dominant temporal change from residual dynamics. Training uses self-supervised real video plus weak scene-dynamics supervision from synthetic Kubric data. They evaluate it on a new ProbeMotion suite covering camera, object, and combined motion. SDM beats CLS and average-pooled backbone baselines and compares favorably with stronger supervised representations on several probes. HF Daily Papers' note
The authors introduce a Structured Dynamics Model that predicts future features while splitting dominant temporal change from residual dynamics. Training uses self-supervised real video plus weak scene-dynamics supervision from synthetic Kubric data. They evaluate it on a new ProbeMotion suite covering camera, object, and combined motion. SDM beats CLS and average-pooled backbone baselines and compares favorably with stronger supervised representations on several probes. HF Daily Papers' note
score 4