Megadose Built for builders and researchers.

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

· HF Daily Papers ·
Accuracy barely moved, but the models’ reasoning language changed sharply.

The paper says base models answered Greek questions while producing zero Greek reasoning traces in 1,000 samples. After SFT, released checkpoints reasoned in the question’s language on about 98% of items, with better grammaticality and little loss in general ability. RL with verifiable rewards then fixed format fallback and reasoning-channel leakage, while an accuracy-only signal did not preserve the same behavioral gains. HF Daily Papers' note

score 5

Categories: Research