Megadose AI progress, ranked and analyzed.

Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States

· HF Daily Papers ·
Mind2Dialogue trains assistants on simulated hidden user beliefs and goals, then tests whether that improves personalized reasoning.

The paper argues that assistant datasets lack supervision grounded in what users privately believe, want, or intend. Its framework simulates evolving mental states, uses an Oracle assistant to respond with access to those states, and distills that behavior into models that do not see the states at deployment. On the reported benchmarks, the full corpus improves personalization metrics across Qwen, Llama, and OLMo baselines, including 26.6 to 40.9 point gains in preference-following generation. The authors also report gains in belief and action reasoning for Qwen and Llama.

HF Daily Papers' note

score 4

Categories: Research