LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
LightNav-0 uses a pretrained VLM as the shared backbone for multiple embodied navigation tasks without task-specific heads.
The paper says the model turns navigation goals into a unified token interface, using dual-channel pointing for spatial intent and an action tokenizer for embodiment-specific trajectories. Its training corpus covers more than 2,000 scenes and 4,000 hours of navigation data. The authors report state-of-the-art monocular success rates across 10 public simulation settings, plus zero-shot real-world tests across robots, scenes, and static or moving targets. HF Daily Papers' note
The paper says the model turns navigation goals into a unified token interface, using dual-channel pointing for spatial intent and an action tokenizer for embodiment-specific trajectories. Its training corpus covers more than 2,000 scenes and 4,000 hours of navigation data. The authors report state-of-the-art monocular success rates across 10 public simulation settings, plus zero-shot real-world tests across robots, scenes, and static or moving targets. HF Daily Papers' note
score 5