Megadose Built for builders and researchers.

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

· ArXiv · AI/CL/LG ·
DECOWAM separates moving-camera, base, and arm effects so a legged robot can predict and control whole-body manipulation more efficiently.

The paper adapts a FastWAM backbone with residual adapters and factorized conditioning for base and arm behavior. It also introduces ARMDOG, a real-robot dataset pairing video, whole-body state/action, and language. In fixed replay, DECOWAM beat FastWAM on future-video and action prediction, including a 21.7% action-MSE reduction with 25.95M trainable adaptation parameters. In 79 closed-loop trials per method, it showed the strongest observed whole-body coordination and base-displacement robustness, while task completion stayed comparable to the best baseline. ArXiv · AI/CL/LG's note

score 5

Categories: Research