HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness
HarnessVLN reports zero-shot navigation gains by checking an MLLM planner’s moves against spatial evidence before acting.
The paper describes an “Agent Harness” that ties together perception, retrieval, grounding, navigation, recovery, and stopping through one tool interface. It keeps event memory and a spatiotemporal graph so the agent can track progress, reuse evidence, and mark failed paths. The same protocol is used for instruction-following and object-goal navigation, with a replaceable executor translating targets into motion. Reported success rates beat prior training-free results on R2R, RxR, HM3D-v2, and HM3D-OVON, and the authors also report humanoid deployment in real-world settings. HF Daily Papers' note
The paper describes an “Agent Harness” that ties together perception, retrieval, grounding, navigation, recovery, and stopping through one tool interface. It keeps event memory and a spatiotemporal graph so the agent can track progress, reuse evidence, and mark failed paths. The same protocol is used for instruction-following and object-goal navigation, with a replaceable executor translating targets into motion. Reported success rates beat prior training-free results on R2R, RxR, HM3D-v2, and HM3D-OVON, and the authors also report humanoid deployment in real-world settings. HF Daily Papers' note
score 5