Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems
Bigger model pools did not reliably improve multi-agent systems, and often made them worse.
The paper evaluates eight ways to choose models for multi-agent systems across routing, majority-vote, and LLM-judge setups on scientific benchmarks. It finds a gap between oracle results and real performance: adding more candidates can fall below the best single base model. The strongest relative gains came from selecting candidates within one model family, rather than mixing arbitrary heterogeneous models. ArXiv · AI/CL/LG's note
The paper evaluates eight ways to choose models for multi-agent systems across routing, majority-vote, and LLM-judge setups on scientific benchmarks. It finds a gap between oracle results and real performance: adding more candidates can fall below the best single base model. The strongest relative gains came from selecting candidates within one model family, rather than mixing arbitrary heterogeneous models. ArXiv · AI/CL/LG's note
score 5