Megadose AI progress, ranked and analyzed.

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

· ArXiv · AI/CL/LG ·
Messier standardizes nearly one million agent-evaluation records so benchmark results can be compared without rerunning them.

The corpus spans 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. The authors add five-agent runs in six less-covered professional and scientific domains, including a recent legal benchmark. Their analysis says function-calling benchmarks are saturated, programming is improving fastest, and enterprise workflows remain hardest. They also find strict all-pass scoring can hide progress and change agent rankings. ArXiv · AI/CL/LG's note

score 6

Categories: Research