FairRSFM: A Biome-Aware Benchmark and Debiasing Framework for Remote Sensing Foundation Models
Aggregate scores are overstating how well remote-sensing foundation models hold up across biomes.
FairRSFM groups 14 terrestrial biome classes into six macro-groups and tests RSFMs under a frozen-backbone protocol. The paper reports gaps that are hidden by headline metrics, including Prithvi-EO-2.0 scoring 90.98% overall macro-F1 on m-EuroSAT but 83.72% on mean worst-group performance. On m-SA-Crop-Type, overall mIoU falls from 27.30% to 18.47% for the Xeric and Mineralogical group. The authors also test mitigation baselines, with results varying by model and task. HF Daily Papers' note
FairRSFM groups 14 terrestrial biome classes into six macro-groups and tests RSFMs under a frozen-backbone protocol. The paper reports gaps that are hidden by headline metrics, including Prithvi-EO-2.0 scoring 90.98% overall macro-F1 on m-EuroSAT but 83.72% on mean worst-group performance. On m-SA-Crop-Type, overall mIoU falls from 27.30% to 18.47% for the Xeric and Mineralogical group. The authors also test mitigation baselines, with results varying by model and task. HF Daily Papers' note
score 4