Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
Accuracy barely moved, but the models’ reasoning language changed sharply.
The paper says base models answered Greek questions while producing zero Greek reasoning traces in 1,000 samples. After SFT, released checkpoints reasoned in the question’s language on about 98% of items, with better grammaticality and little loss in general ability. RL with verifiable rewards then fixed format fallback and reasoning-channel leakage, while an accuracy-only signal did not preserve the same behavioral gains. HF Daily Papers' note
The paper says base models answered Greek questions while producing zero Greek reasoning traces in 1,000 samples. After SFT, released checkpoints reasoned in the question’s language on about 98% of items, with better grammaticality and little loss in general ability. RL with verifiable rewards then fixed format fallback and reasoning-channel leakage, while an accuracy-only signal did not preserve the same behavioral gains. HF Daily Papers' note
score 5