Megadose Built for builders and researchers.

MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability

· ArXiv · AI/CL/LG ·
MASBench tests collaboration mechanics when each agent sees only part of the problem.

The paper introduces a benchmark for LLM-based multi-agent systems under partially observable constraints. It groups tasks into Reasoning, Scheduling, and Game categories to evaluate protocol, memory, and routing mechanisms. Its metrics cover task performance, communication cost, and cost effectiveness, aiming to measure both outcomes and overhead. Code is listed as available. ArXiv · AI/CL/LG's note

score 5

Categories: Research