The Illusion of Cross-Lingual Safety in Low-Resource Languages
Harmful prompts in the tested African languages carried under 10% of the English refusal signal in most model-language pairs.
The paper tests Twi, Hausa, Amharic, and Swahili with LoDNA, a dataset pairing literal translations with culturally localized prompts. The authors probe hidden-state refusal representations rather than relying only on generated outputs. They find that models often preserve semantic meaning while failing to route those prompts into safety mechanisms. ArXiv · AI/CL/LG's note
The paper tests Twi, Hausa, Amharic, and Swahili with LoDNA, a dataset pairing literal translations with culturally localized prompts. The authors probe hidden-state refusal representations rather than relying only on generated outputs. They find that models often preserve semantic meaning while failing to route those prompts into safety mechanisms. ArXiv · AI/CL/LG's note
score 5