On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens
The paper tests LLM translation on 500 Chinese-Japanese segments from *Dream of the Red Chamber* and finds culture is still a weak point.
The authors group the dataset across cultural categories and evaluate frontier LLM-based MT systems on culturally loaded expressions. They report three problem areas: model performance gaps, disagreement among human evaluators with different backgrounds, and automatic metrics that do not reliably judge this task. The claim is narrower than general MT quality: surface fluency can miss meanings embedded in social and cultural context. ArXiv · AI/CL/LG's note
The authors group the dataset across cultural categories and evaluate frontier LLM-based MT systems on culturally loaded expressions. They report three problem areas: model performance gaps, disagreement among human evaluators with different backgrounds, and automatic metrics that do not reliably judge this task. The claim is narrower than general MT quality: surface fluency can miss meanings embedded in social and cultural context. ArXiv · AI/CL/LG's note
score 4