Megadose AI progress, ranked and analyzed.

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

· HF Daily Papers ·
The benchmark’s best-tested model still solved only 52.0% of the long collaboration tasks.

AgentWorld tests 3-20 LLM agents over 50-plus interaction rounds inside an MMORPG-style sandbox, with agents forced to coordinate from separate black-box views. The tasks require communication, planning, and resource sharing across asymmetric roles. The paper adds Causal Collaboration Effectiveness, a metric for measuring how much of the team’s work actually helped the outcome. Reported failures include communication breakdowns, role confusion, and loss of shared plans across rounds. HF Daily Papers' note

score 5

Categories: Research