MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation
The paper finds that memory recall scores can look strong while the assistant almost never uses those memories naturally.
In a 4-month deployment with 40 users and 1,872 sessions, Direct QA accuracy ranged from 19.7% to 70.1% across seven memory conditions, but user satisfaction did not move. The authors argue that direct recall and conversational usefulness are measuring different things. Their MemUse benchmark uses real user-cued moments from the deployment to judge whether prior context is integrated naturally. With the model and context held fixed, a system scoring 78.8% on Direct QA referenced only 7.9% of those facts in conversation. HF Daily Papers' note
In a 4-month deployment with 40 users and 1,872 sessions, Direct QA accuracy ranged from 19.7% to 70.1% across seven memory conditions, but user satisfaction did not move. The authors argue that direct recall and conversational usefulness are measuring different things. Their MemUse benchmark uses real user-cued moments from the deployment to judge whether prior context is integrated naturally. With the model and context held fixed, a system scoring 78.8% on Direct QA referenced only 7.9% of those facts in conversation. HF Daily Papers' note
score 4