Testing Interchangeability in LLM Agent Teams
Swapping role-matched agents barely hurt scores, but made teams spend much more communication to get work done.
The paper tests multi-agent teams formed from the same base model, then trades agents between teams on held-out tasks. Compared with a placebo roster disruption, swaps raised communication per unit of progress by 16 to 63 percent. In Hanabi, a swapped agent was costlier than a new one, which the authors tie to conventions learned with its old partner. Longer team histories increased the penalty; greedy decoding reduced it. ArXiv · AI/CL/LG's note
The paper tests multi-agent teams formed from the same base model, then trades agents between teams on held-out tasks. Compared with a placebo roster disruption, swaps raised communication per unit of progress by 16 to 63 percent. In Hanabi, a swapped agent was costlier than a new one, which the authors tie to conventions learned with its old partner. Longer team histories increased the penalty; greedy decoding reduced it. ArXiv · AI/CL/LG's note
score 5