PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems
The paper argues agent training is still built around the wrong shape of user behavior.
The authors say their analysis of 16,000 real-world sessions found 75.9% were multi-turn, while many datasets and benchmarks still center on complete single-turn requests. PersonaForge simulates user-agent exchanges using persona variables, behavior controls calibrated to real-user statistics, and seed-query reconstruction. The team also built a 6.3K-record training set and a 138-task benchmark across more than 20 professional domains. In Qwen3.5-27B experiments, PersonaForge training raised the composite score by 4.1%, with the biggest gains in task completion and response quality. ArXiv · AI/CL/LG's note
The authors say their analysis of 16,000 real-world sessions found 75.9% were multi-turn, while many datasets and benchmarks still center on complete single-turn requests. PersonaForge simulates user-agent exchanges using persona variables, behavior controls calibrated to real-user statistics, and seed-query reconstruction. The team also built a 6.3K-record training set and a 138-task benchmark across more than 20 professional domains. In Qwen3.5-27B experiments, PersonaForge training raised the composite score by 4.1%, with the biggest gains in task completion and response quality. ArXiv · AI/CL/LG's note
score 5