Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
The paper trains assistants on simulated hidden user beliefs and goals, then removes that privileged access at deployment.
Mind2Dialogue uses a psychology-guided simulator to keep user characteristics consistent while mental states evolve through conversation. An Oracle assistant answers with access to those simulated states, and smaller models are distilled from those responses. The authors report gains over Qwen, Llama, and OLMo instruction-tuned baselines, including 26.6 to 40.9 percentage-point improvements in preference-following generation. ArXiv · AI/CL/LG's note
Mind2Dialogue uses a psychology-guided simulator to keep user characteristics consistent while mental states evolve through conversation. An Oracle assistant answers with access to those simulated states, and smaller models are distilled from those responses. The authors report gains over Qwen, Llama, and OLMo instruction-tuned baselines, including 26.6 to 40.9 percentage-point improvements in preference-following generation. ArXiv · AI/CL/LG's note
score 4