Megadose Built for builders and researchers.

Long-WAM: Scaling the Context of World-Action Models

· ArXiv · AI/CL/LG ·
Long-WAM’s gain comes from using longer visual history under real-time robot-control limits, especially with autoregressive pretraining.

The paper reports RoboCasa GR-1 success rising from 63.3% to 78.7% when context expands from 0.0 to 19.2 seconds. A bidirectionally pretrained initialization did not show the same net gain, while robot-domain autoregressive pretraining improved peak results on GR-1 and LIBERO-Long. The system keeps future-video latent prediction while running in real time, with an RTX 5090 action chunk taking 107.4 ms. On Unitree G1 and YAM robots, it handled dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking where two compared methods failed all 20 trials. ArXiv · AI/CL/LG's note

score 6

Categories: Research