EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards
EarthVerse tests agents on 405 reproducible natural-hazard investigations, and even the top systems struggle with strict end-to-end reliability.
The benchmark is built from 199 documented events across 19 hazard families. Agents must work through packaged evidence, select compatible sources, run transparent calculations, reconcile differences, and preserve provenance. In the reported evaluation of 25 model and agent systems, the best mean answer-unit accuracy was 84.65%, but the highest Strict@95 score was 34.81%. The authors say that gap shows agents can finish individual steps while losing consistency across evidence, scale, units, calculations, and physical interpretation. ArXiv · AI/CL/LG's note
The benchmark is built from 199 documented events across 19 hazard families. Agents must work through packaged evidence, select compatible sources, run transparent calculations, reconcile differences, and preserve provenance. In the reported evaluation of 25 model and agent systems, the best mean answer-unit accuracy was 84.65%, but the highest Strict@95 score was 34.81%. The authors say that gap shows agents can finish individual steps while losing consistency across evidence, scale, units, calculations, and physical interpretation. ArXiv · AI/CL/LG's note
score 5