Megadose AI progress, ranked and analyzed.

When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs

· ArXiv · AI/CL/LG ·
Open-weight LLMs still miss ordinary culture, especially across languages.

The paper introduces TriviaRoomQA, a quiz-style benchmark with 3,300 parallel multiple-choice questions in six European languages, plus 5,340 French-only questions. It tests 30 open-weight models from 7B to 70B parameters across 288 topics. The models do better on history, geography, and mathematics than on celebrities, music, movies, and news. Performance also shifts by language on the same questions, suggesting factual access is not fully language-independent. ArXiv · AI/CL/LG's note

score 4

Categories: Research