Skill Issue: Are Skills Language-Invariant in LLMs?
The same model’s game-playing skill changes when only the language interface changes.
The paper tests this with multilingual self-play, holding the model, rules, opponent, state space, and actions fixed. Across three open-weight models, eight languages, and six TextArena games, the authors report different win-loss margins, invalid-action rates, and strategies by language. Some failures appear in spatial reasoning, card-based choices, and optimal move selection. In some cases, switching only the intermediate reasoning language recovered much of the lost performance. HF Daily Papers' note
The paper tests this with multilingual self-play, holding the model, rules, opponent, state space, and actions fixed. Across three open-weight models, eight languages, and six TextArena games, the authors report different win-loss margins, invalid-action rates, and strategies by language. Some failures appear in spatial reasoning, card-based choices, and optimal move selection. In some cases, switching only the intermediate reasoning language recovered much of the lost performance. HF Daily Papers' note
score 5