CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
CultureConverse tests whether assistants can handle culturally grounded help across multi-turn dialogue, not just answer culture trivia.
The harness covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains. Its dataset includes 14,610 benchmark episodes and 274,295 oracle-guided dialogues. In tests across 18 models, GPT-5 mini scored highest on assistance quality. The authors say human annotation supports their evaluation as a proxy for human judgment, and that fine-tuning on selected samples improved both in-domain help and transfer to other cultural and safety benchmarks. ArXiv · AI/CL/LG's note
The harness covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains. Its dataset includes 14,610 benchmark episodes and 274,295 oracle-guided dialogues. In tests across 18 models, GPT-5 mini scored highest on assistance quality. The authors say human annotation supports their evaluation as a proxy for human judgment, and that fine-tuning on selected samples improved both in-domain help and transfer to other cultural and safety benchmarks. ArXiv · AI/CL/LG's note
score 4