SuperNav: An Agentic Navigation System for Any Task in Any Scene
SuperNav keeps the language model as a planner and decision-maker, while separate navigation tools handle motion.
The paper describes a pretrained multimodal LLM placed inside an agent harness, without navigation-specific fine-tuning. It uses navigation skills, physical-interaction tools, context tracking, and a visual-point interface that lets the model choose destinations in images and revise choices from execution feedback. The authors report stronger results than four baselines across instance-level, multi-object, and demand-driven navigation tasks, plus category-level tests on HM3D and a real quadruped deployment. HF Daily Papers' note
The paper describes a pretrained multimodal LLM placed inside an agent harness, without navigation-specific fine-tuning. It uses navigation skills, physical-interaction tools, context tracking, and a visual-point interface that lets the model choose destinations in images and revise choices from execution feedback. The authors report stronger results than four baselines across instance-level, multi-object, and demand-driven navigation tasks, plus category-level tests on HM3D and a real quadruped deployment. HF Daily Papers' note
score 5