OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
The benchmark tests whether open-vocabulary EO models can handle broader geospatial categories and more varied language queries.
OVEarth-Bench adds hierarchical category coverage, positive and negative expressions, and vocabulary, referring, and reasoning queries. It evaluates mask and box localization under a unified zero-shot protocol. The authors report that current methods remain limited, with MLLM-based systems performing best overall. EO-specific methods generally trail general models in the evaluation. HF Daily Papers' note
OVEarth-Bench adds hierarchical category coverage, positive and negative expressions, and vocabulary, referring, and reasoning queries. It evaluates mask and box localization under a unified zero-shot protocol. The authors report that current methods remain limited, with MLLM-based systems performing best overall. EO-specific methods generally trail general models in the evaluation. HF Daily Papers' note
score 4