ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
A demonstration clip becomes a reusable action control signal for new video environments.
ShadowDancer builds “shadow pairs”: videos with the same dynamics but independently resampled appearance. It trains by predicting one shadow from the other, forcing the model to keep motion and discard appearance. The authors say this enables frame-level control without action labels, motion estimators, or fine-tuning. They report better action transfer and long rollouts than baseline systems, including an 86% average blinded win rate in rollout comparisons. HF Daily Papers' note
ShadowDancer builds “shadow pairs”: videos with the same dynamics but independently resampled appearance. It trains by predicting one shadow from the other, forcing the model to keep motion and discard appearance. The authors say this enables frame-level control without action labels, motion estimators, or fine-tuning. They report better action transfer and long rollouts than baseline systems, including an 86% average blinded win rate in rollout comparisons. HF Daily Papers' note
score 5