When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs
Open-weight LLMs still miss ordinary culture, especially across languages.
The paper introduces TriviaRoomQA, a quiz-style benchmark with 3,300 parallel multiple-choice questions in six European languages, plus 5,340 French-only questions. It tests 30 open-weight models from 7B to 70B parameters across 288 topics. The models do better on history, geography, and mathematics than on celebrities, music, movies, and news. Performance also shifts by language on the same questions, suggesting factual access is not fully language-independent. ArXiv · AI/CL/LG's note
The paper introduces TriviaRoomQA, a quiz-style benchmark with 3,300 parallel multiple-choice questions in six European languages, plus 5,340 French-only questions. It tests 30 open-weight models from 7B to 70B parameters across 288 topics. The models do better on history, geography, and mathematics than on celebrities, music, movies, and news. Performance also shifts by language on the same questions, suggesting factual access is not fully language-independent. ArXiv · AI/CL/LG's note
score 4