Megadose Built for builders and researchers.

Long-WAM: Scaling the Context of World-Action Models

· HF Daily Papers ·
Long-WAM’s key result is that longer robot video history helps only when the model is trained to predict forward causally.

The paper reports RoboCasa GR-1 success rising from 63.3% to 78.7% when context expands to 19.2 seconds, while a bidirectionally pretrained version shows no net gain. Its AR pretraining uses robot and egocentric video without action labels, then keeps that history-to-future structure during adaptation. The system is also built for real-time deployment, with RTX 5090 action chunks taking 107.4 ms including future-video latent prediction. On physical robots, it reports 95% success on dynamic cup stacking, versus zero successes in 20 trials for Pi0.5 and Fast-WAM.

HF Daily Papers' note

score 5

Categories: Research