Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
TAMP-Nav lets a VLM choose pixels, then hands the movement to a SLAM controller.
The paper frames that as a better fit for vision-language models’ 2D training than forcing them into navigation action spaces. It adds selective reasoning and memory so the agent reasons and stores detailed trajectory state only at critical points. A two-level alignment setup uses GRPO with outcome and process rewards to tie planning to environmental feedback. The authors report state-of-the-art results, including 66.2% success rate on R2R-CE with 90k training trajectories. HF Daily Papers' note
The paper frames that as a better fit for vision-language models’ 2D training than forcing them into navigation action spaces. It adds selective reasoning and memory so the agent reasons and stores detailed trajectory state only at critical points. A two-level alignment setup uses GRPO with outcome and process rewards to tie planning to environmental feedback. The authors report state-of-the-art results, including 66.2% success rate on R2R-CE with 90k training trajectories. HF Daily Papers' note
score 4