SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery
SatNav builds 118K city-scale UAV navigation episodes from satellite imagery, then shows current VLN agents still struggle at that range.
The benchmark spans 59 scenes in 18 cities, with average routes of 379 meters. It uses satellite crops as a stand-in for nadir UAV views and defines Boundary, Landmark, and Route task families to test memory and geospatial grounding. The authors also introduce SwiftVLN for memory ablations. Their transfer tests indicate models trained on satellite imagery can run on real-flight UAV observations. HF Daily Papers' note
The benchmark spans 59 scenes in 18 cities, with average routes of 379 meters. It uses satellite crops as a stand-in for nadir UAV views and defines Boundary, Landmark, and Route task families to test memory and geospatial grounding. The authors also introduce SwiftVLN for memory ablations. Their transfer tests indicate models trained on satellite imagery can run on real-flight UAV observations. HF Daily Papers' note
score 4