Megadose AI progress, ranked daily.

MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

· HF Daily Papers ·
The paper finds that memory recall scores can look strong while the assistant almost never uses those memories naturally.

In a 4-month deployment with 40 users and 1,872 sessions, Direct QA accuracy ranged from 19.7% to 70.1% across seven memory conditions, but user satisfaction did not move. The authors argue that direct recall and conversational usefulness are measuring different things. Their MemUse benchmark uses real user-cued moments from the deployment to judge whether prior context is integrated naturally. With the model and context held fixed, a system scoring 78.8% on Direct QA referenced only 7.9% of those facts in conversation. HF Daily Papers' note

score 4

Categories: Research