Multilingual GSM-Symbolic: What determines capability transfer across languages?
A new benchmark ties cross-language math transfer mostly to model size, language resources, reasoning strength, and typological distance.
The paper introduces Multilingual GSM-Symbolic, with 30,000 item-matched math QA pairs across 15 languages. Its analysis says model size is the largest determinant of capability transfer, followed by language resource level and reasoning, while typological distance hurts transfer. The authors estimate that a 32B model on Marathi performs like a 10B model on English. Their framework explains most between-language variation and can predict performance on an unseen language within 6.0 percentage points. HF Daily Papers' note
The paper introduces Multilingual GSM-Symbolic, with 30,000 item-matched math QA pairs across 15 languages. Its analysis says model size is the largest determinant of capability transfer, followed by language resource level and reasoning, while typological distance hurts transfer. The authors estimate that a 32B model on Marathi performs like a 10B model on English. Their framework explains most between-language variation and can predict performance on an unseen language within 6.0 percentage points. HF Daily Papers' note
score 5