Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Messier standardizes nearly one million agent-evaluation records so benchmark results can be compared without rerunning them.
The corpus spans 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. The authors add five-agent runs in six less-covered professional and scientific domains, including a recent legal benchmark. Their analysis says function-calling benchmarks are saturated, programming is improving fastest, and enterprise workflows remain hardest. They also find strict all-pass scoring can hide progress and change agent rankings. ArXiv · AI/CL/LG's note
The corpus spans 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. The authors add five-agent runs in six less-covered professional and scientific domains, including a recent legal benchmark. Their analysis says function-calling benchmarks are saturated, programming is improving fastest, and enterprise workflows remain hardest. They also find strict all-pass scoring can hide progress and change agent rankings. ArXiv · AI/CL/LG's note
score 6