Megadose Built for builders and researchers.

You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue

· HF Daily Papers ·
Rejected changes still pull models off course.

The paper introduces Intent-Eval, a benchmark for testing whether models can separate what a user mentioned from what remains in effect. It reports failures across tool actions, code, databases, and math when proposals are rejected or later replaced. The authors call the pattern “mentioned-as-in-effect confusion” and propose Intent-OPSD to train models on active user intent across the full dialogue. HF Daily Papers' note

score 5

Categories: Research