WorldBench: Culturally Grounded Benchmark for Multilingual Agents
Frontier agents topped out at 49.2% on a benchmark built around multilingual everyday workflows.
WorldBench tests 1,600 persona-grounded tasks across seven languages and eight cultures, with agents acting through structured actions in a sandbox. The authors say the tasks were refined with human annotators who had language- and culture-specific expertise. They introduce Constrained Task Success, a metric meant to capture completion while also checking whether the agent preserves the surrounding environment. The reported gap is not just accuracy: models often fail to maintain state, especially on long-horizon tasks. ArXiv · AI/CL/LG's note
WorldBench tests 1,600 persona-grounded tasks across seven languages and eight cultures, with agents acting through structured actions in a sandbox. The authors say the tasks were refined with human annotators who had language- and culture-specific expertise. They introduce Constrained Task Success, a metric meant to capture completion while also checking whether the agent preserves the surrounding environment. The reported gap is not just accuracy: models often fail to maintain state, especially on long-horizon tasks. ArXiv · AI/CL/LG's note
score 5