MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability
MASBench tests collaboration mechanics when each agent sees only part of the problem.
The paper introduces a benchmark for LLM-based multi-agent systems under partially observable constraints. It groups tasks into Reasoning, Scheduling, and Game categories to evaluate protocol, memory, and routing mechanisms. Its metrics cover task performance, communication cost, and cost effectiveness, aiming to measure both outcomes and overhead. Code is listed as available. ArXiv · AI/CL/LG's note
The paper introduces a benchmark for LLM-based multi-agent systems under partially observable constraints. It groups tasks into Reasoning, Scheduling, and Game categories to evaluate protocol, memory, and routing mechanisms. Its metrics cover task performance, communication cost, and cost effectiveness, aiming to measure both outcomes and overhead. Code is listed as available. ArXiv · AI/CL/LG's note
score 5