Megadose AI progress, ranked and analyzed.

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

· ArXiv · AI/CL/LG ·
Frontier agents topped out at 49.2% on a benchmark built around multilingual everyday workflows.

WorldBench tests 1,600 persona-grounded tasks across seven languages and eight cultures, with agents acting through structured actions in a sandbox. The authors say the tasks were refined with human annotators who had language- and culture-specific expertise. They introduce Constrained Task Success, a metric meant to capture completion while also checking whether the agent preserves the surrounding environment. The reported gap is not just accuracy: models often fail to maintain state, especially on long-horizon tasks. ArXiv · AI/CL/LG's note

score 5

Categories: Research