One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Training against a single simulated user can teach an agent to exploit that simulator, then fail elsewhere.
The paper names the failure “simulator collapse”: a mode-collapsed LLM simulator pushes policies toward narrow strategies that do not transfer well to new simulators or real users. The authors propose Verbalized Sampling at inference time and Co-Training against a population of trainable simulators during training. Across three multi-turn benchmarks, they report held-out success gains of up to 9% and 14%, respectively, with a human study showing a similar real-user gain. They also release SCOPE, an open-source framework for population co-training in multi-agent RL. ArXiv · AI/CL/LG's note
The paper names the failure “simulator collapse”: a mode-collapsed LLM simulator pushes policies toward narrow strategies that do not transfer well to new simulators or real users. The authors propose Verbalized Sampling at inference time and Co-Training against a population of trainable simulators during training. Across three multi-turn benchmarks, they report held-out success gains of up to 9% and 14%, respectively, with a human study showing a similar real-user gain. They also release SCOPE, an open-source framework for population co-training in multi-agent RL. ArXiv · AI/CL/LG's note
score 5