TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
TANGO maps language and egocentric RGB directly to 29-DoF whole-body humanoid actions for moving through clutter.
The paper frames cluttered indoor navigation as a 3D whole-body control problem, not just 2D path planning. Its training data is generated entirely in simulation, combining path planning, whole-body motion generation, obstacle-aware editing, and RL tracking. In simulation, the authors report state-of-the-art vision-language navigation results and stronger performance than modular baselines on obstacle-negotiation scenes. They also report zero-shot deployment on a Unitree G1 in real cluttered scenes, without real-world navigation training data. ArXiv · AI/CL/LG's note
The paper frames cluttered indoor navigation as a 3D whole-body control problem, not just 2D path planning. Its training data is generated entirely in simulation, combining path planning, whole-body motion generation, obstacle-aware editing, and RL tracking. In simulation, the authors report state-of-the-art vision-language navigation results and stronger performance than modular baselines on obstacle-negotiation scenes. They also report zero-shot deployment on a Unitree G1 in real cluttered scenes, without real-world navigation training data. ArXiv · AI/CL/LG's note
score 5