Megadose Built for builders and researchers.

When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents

· HF Daily Papers ·
The paper isolates “intent drift” as a testable agent failure: old user instructions keep affecting the final action after the user has changed course.

The authors introduce IntentFlux, a benchmark that turns verifiable tasks into multi-turn dialogues with controlled changes in user intent. In their calibration, mean task score drops as superseded or withdrawn information accumulates. Across eight models, agents perform worse when recovering the same final task from an evolving dialogue than when given it directly. Their StateForge method improves results by tracking active requirements before generation, but even ground-truth final state does not close the gap. HF Daily Papers' note

score 5

Categories: Research