Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Cultivar tests whether translation models handle local context, not just language pairs.
The paper introduces a localized subset of FLORES for source-contrastive translation evaluation. Paired with unlocalized versions, it is meant to expose possible benchmark contamination and weak localization robustness. The authors benchmark 32 open-weight models and report that MT-specialized models are less robust, some models may be overfit to FLORES, and US-centered content is translated better than content from other locales. ArXiv · AI/CL/LG's note
The paper introduces a localized subset of FLORES for source-contrastive translation evaluation. Paired with unlocalized versions, it is meant to expose possible benchmark contamination and weak localization robustness. The authors benchmark 32 open-weight models and report that MT-specialized models are less robust, some models may be overfit to FLORES, and US-centered content is translated better than content from other locales. ArXiv · AI/CL/LG's note
score 5