InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data
The paper presents InternW0-Δ as a unified robot world-action model trained on more than 20,000 hours of open processed data.
It combines video dynamics, semantic guidance from a frozen VLM, 4D geometry and motion priors, and action generation inside a Mixture-of-Transformers framework. The authors say a “Causal Imprint” component gives the action expert predictive scene-change representations without rolling out future video at inference. The training corpus mixes robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data under a shared state-action format. They report stronger results than prior methods across simulation benchmarks and real-robot platforms, and say code, weights, infrastructure, pipeline, and permitted processed data will be open sourced. HF Daily Papers' note
It combines video dynamics, semantic guidance from a frozen VLM, 4D geometry and motion priors, and action generation inside a Mixture-of-Transformers framework. The authors say a “Causal Imprint” component gives the action expert predictive scene-change representations without rolling out future video at inference. The training corpus mixes robot demonstrations, UMI data, egocentric human demonstrations, and Ego2Robot data under a shared state-action format. They report stronger results than prior methods across simulation benchmarks and real-robot platforms, and say code, weights, infrastructure, pipeline, and permitted processed data will be open sourced. HF Daily Papers' note
score 6