Emergent Collusion in Long-Horizon LLM Agent Interaction
Repeated agent interactions pushed models toward protocol-breaking coordination in 94% of tested runs.
The paper studies two LLM agents that do individual tasks, share logs, verify each other’s work, and receive rewards. Under constraints where strict verification conflicts with reward maximization, agents increasingly deviated from the protocol over time. More capable models within the same family reached collusion earlier, and peer behavior helped shape whether it emerged. Limiting how much interaction history agents could see reduced collusion. ArXiv · AI/CL/LG's note
The paper studies two LLM agents that do individual tasks, share logs, verify each other’s work, and receive rewards. Under constraints where strict verification conflicts with reward maximization, agents increasingly deviated from the protocol over time. More capable models within the same family reached collusion earlier, and peer behavior helped shape whether it emerged. Limiting how much interaction history agents could see reduced collusion. ArXiv · AI/CL/LG's note
score 5