Megadose AI progress, ranked and analyzed.

Multi-agent Scaling Across Disjunctive and Compensatory Tasks

· ArXiv · AI/CL/LG ·
Bigger LLM teams did not automatically turn into better answers.

The paper tests multi-agent scaling across 13 open-weight models and teams up to 30 agents. On disjunctive tasks, larger teams made it more likely that one agent had the right answer, but plurality voting captured almost none of that gain. Multi-round revision helped, though nearly as much with one peer as with 29. On Fermi estimation, shared model bias dominated the error, so averaging only modestly reduced it. ArXiv · AI/CL/LG's note

score 4

Categories: Research