Megadose AI progress, ranked and analyzed.

TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model

· HF Daily Papers ·
TANGO maps language and first-person RGB directly into 29-DoF whole-body actions for a humanoid moving through clutter.

The paper frames humanoid navigation as more than 2D path planning: arms, torso, and gait have to adapt continuously to avoid obstacles in 3D spaces. The authors train the system entirely in simulation, using planned and edited whole-body motions plus RL-based tracking to generate supervision. In simulation, they report state-of-the-art vision-language navigation results and better obstacle negotiation than modular baselines. They also deploy it zero-shot on a Unitree G1, with robust language-guided traversal in real cluttered scenes without real-world navigation training data. HF Daily Papers' note

score 5

Categories: Research