AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
The benchmark’s best-tested model still solved only 52.0% of the long collaboration tasks.
AgentWorld tests 3-20 LLM agents over 50-plus interaction rounds inside an MMORPG-style sandbox, with agents forced to coordinate from separate black-box views. The tasks require communication, planning, and resource sharing across asymmetric roles. The paper adds Causal Collaboration Effectiveness, a metric for measuring how much of the team’s work actually helped the outcome. Reported failures include communication breakdowns, role confusion, and loss of shared plans across rounds. HF Daily Papers' note
AgentWorld tests 3-20 LLM agents over 50-plus interaction rounds inside an MMORPG-style sandbox, with agents forced to coordinate from separate black-box views. The tasks require communication, planning, and resource sharing across asymmetric roles. The paper adds Causal Collaboration Effectiveness, a metric for measuring how much of the team’s work actually helped the outcome. Reported failures include communication breakdowns, role confusion, and loss of shared plans across rounds. HF Daily Papers' note
score 5