Robostral Navigate
An 8B vision-language navigation model claims state-of-the-art results using only a single RGB camera.
Robostral Navigate predicts waypoints directly in image space, avoiding depth sensors, multi-camera rigs, pre-built maps, and robot-specific coordinates. The authors say this makes it portable across wheeled, legged, and aerial robots without recalibration. They trained with 2.4 million simulated trajectories across 350,000 scenes, plus a prefix-caching recipe that cut training tokens by 22x. On R2R-CE it reports a 77.4% success rate; on RxR-CE, 75.1%.
HF Daily Papers' note
Robostral Navigate predicts waypoints directly in image space, avoiding depth sensors, multi-camera rigs, pre-built maps, and robot-specific coordinates. The authors say this makes it portable across wheeled, legged, and aerial robots without recalibration. They trained with 2.4 million simulated trajectories across 350,000 scenes, plus a prefix-caching recipe that cut training tokens by 22x. On R2R-CE it reports a 77.4% success rate; on RxR-CE, 75.1%.
HF Daily Papers' note
score 6