Long-WAM: Scaling the Context of World-Action Models
Long-WAM’s key result is that longer robot video history helps only when the model is trained to predict forward causally.
The paper reports RoboCasa GR-1 success rising from 63.3% to 78.7% when context expands to 19.2 seconds, while a bidirectionally pretrained version shows no net gain. Its AR pretraining uses robot and egocentric video without action labels, then keeps that history-to-future structure during adaptation. The system is also built for real-time deployment, with RTX 5090 action chunks taking 107.4 ms including future-video latent prediction. On physical robots, it reports 95% success on dynamic cup stacking, versus zero successes in 20 trials for Pi0.5 and Fast-WAM.
HF Daily Papers' note
The paper reports RoboCasa GR-1 success rising from 63.3% to 78.7% when context expands to 19.2 seconds, while a bidirectionally pretrained version shows no net gain. Its AR pretraining uses robot and egocentric video without action labels, then keeps that history-to-future structure during adaptation. The system is also built for real-time deployment, with RTX 5090 action chunks taking 107.4 ms including future-video latent prediction. On physical robots, it reports 95% success on dynamic cup stacking, versus zero successes in 20 trials for Pi0.5 and Fast-WAM.
HF Daily Papers' note
score 5