PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation
Panoramic views helped navigation only after the model was changed to plan farther, handle branching routes, and read spatial layout across directions.
PanoVLN uses panoramas for vision-and-language navigation, but the paper says swapping in wider images alone produced limited gains. The authors add longer action-sequence prediction with confidence-guided execution, build training routes with more branching decisions, and combine semantic and geometric RGB panorama features without adding visual tokens. With a 4B backbone and RGB-only input, it reports success-rate gains of 11.9% on R2R-CE Val-Unseen and 8.7% on RxR-CE Val-Unseen over the previous SOTA. Real-world quadruped tests showed faster navigation with fewer pauses than earlier VLN methods. HF Daily Papers' note
PanoVLN uses panoramas for vision-and-language navigation, but the paper says swapping in wider images alone produced limited gains. The authors add longer action-sequence prediction with confidence-guided execution, build training routes with more branching decisions, and combine semantic and geometric RGB panorama features without adding visual tokens. With a 4B backbone and RGB-only input, it reports success-rate gains of 11.9% on R2R-CE Val-Unseen and 8.7% on RxR-CE Val-Unseen over the previous SOTA. Real-world quadruped tests showed faster navigation with fewer pauses than earlier VLN methods. HF Daily Papers' note
score 4