Megadose AI progress, ranked and analyzed.

MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

· ArXiv · AI/CL/LG ·
The benchmark finds LLMs still fail hard at geospatial computation, even when given the right facts.

MultiGlobeQA contains 46,060 multilingual question-answer pairs across 201 countries and territories, built from three knowledge graphs. The paper says models do best on topological relations and directions, but collapse on grid indexing and shape computation. Retrieval and tool use improve results, yet performance remains below two thirds even with gold facts supplied. The authors also report weaker performance on low-income regions, with gold facts widening that gap rather than closing it. ArXiv · AI/CL/LG's note

score 5

Categories: Research