Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning
A 3.35B model hit above 93% in-language reasoning across 60 languages by changing the supervised fine-tuning mix.
The paper argues that models can learn to reason in the user’s prompt language without reasoning examples for every target language. Its Tiny Aya L2-Thinker is tested on six benchmarks covering math, commonsense, instruction following, open-ended generation, and cultural reasoning. The authors say broader language coverage, multilingual non-reasoning data, and a strong English reasoning base were key to generalizing to held-out languages. They release the model weights and multilingual reasoning data for further research. HF Daily Papers' note
The paper argues that models can learn to reason in the user’s prompt language without reasoning examples for every target language. Its Tiny Aya L2-Thinker is tested on six benchmarks covering math, commonsense, instruction following, open-ended generation, and cultural reasoning. The authors say broader language coverage, multilingual non-reasoning data, and a strong English reasoning base were key to generalizing to held-out languages. They release the model weights and multilingual reasoning data for further research. HF Daily Papers' note
score 5