MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
MIMESIS is a 9B user simulator built to make agent training less dependent on costly human interactions.
The paper says the simulator is trained on human conversations, explicit reasoning supervision, and 13 behavior patterns drawn from real users. It reports a SOUL-Index of 65.7, ahead of the strongest frontier-model baseline cited in the abstract. When frozen and used for multi-turn reinforcement learning, MIMESIS-trained agents beat GPT-5.5-trained agents across eight environments and nine unseen user simulators. The authors also introduce Coached On-Policy Self-Distillation, using simulator reasoning traces to turn sparse rewards into denser coaching signals. HF Daily Papers' note
The paper says the simulator is trained on human conversations, explicit reasoning supervision, and 13 behavior patterns drawn from real users. It reports a SOUL-Index of 65.7, ahead of the strongest frontier-model baseline cited in the abstract. When frozen and used for multi-turn reinforcement learning, MIMESIS-trained agents beat GPT-5.5-trained agents across eight environments and nine unseen user simulators. The authors also introduce Coached On-Policy Self-Distillation, using simulator reasoning traces to turn sparse rewards into denser coaching signals. HF Daily Papers' note
score 5