Megadose AI progress, ranked and analyzed.

Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models

· ArXiv · AI/CL/LG ·
In this benchmark, Claude Opus 5 is the clear laggard while five other frontier models cluster near the top.

Peter Potash tests six models in a two-agent game where one model instance asks exactly `log2 N` yes/no questions to find a hidden Wikipedia lead paragraph, while another answers from the target alone. Across 408 games, Claude Opus 5 wins 28 of 68, versus 45 to 56 wins for GLM-5.3, GPT-5.6 Sol, Grok 4.6, Gemini 3.8 Flash, and Kimi K3. For the stronger five, accuracy falls with larger document sets and fits a per-question reliability model of `p=0.928`. The paper says failures split roughly between wrong answers and poor discrimination, with information per question strongly tied to win rate.

ArXiv · AI/CL/LG's note

score 4

Categories: Research