WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
The benchmark tests whether models can forecast matches before kickoff, then grades them after the real result is known.
WorldCupArena uses the 2026 FIFA World Cup as its first run, covering 104 matches and 13 systems. Models either receive a shared evidence package or gather information themselves, then predict results, scores, players, events, statistics, and tournament outcomes. The paper says similar result accuracy can hide wider gaps in detailed predictions, with the best system only slightly ahead of betting-market and fan baselines on winners and exact scores but clearer on scoreline scoring. Code, prompts, predictions, and evaluation scripts are open sourced. ArXiv · AI/CL/LG's note
WorldCupArena uses the 2026 FIFA World Cup as its first run, covering 104 matches and 13 systems. Models either receive a shared evidence package or gather information themselves, then predict results, scores, players, events, statistics, and tournament outcomes. The paper says similar result accuracy can hide wider gaps in detailed predictions, with the best system only slightly ahead of betting-market and fan baselines on winners and exact scores but clearer on scoreline scoring. Code, prompts, predictions, and evaluation scripts are open sourced. ArXiv · AI/CL/LG's note
score 4