Megadose AI progress, ranked and analyzed.

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

· ArXiv · AI/CL/LG ·
The benchmark finds agent memory systems fall off as soon as “knowing the user” requires abstraction, not just recall.

Setoka tests personalized agents across four levels: semantic memory, episodic memory, behavior patterns, and personality traits. The authors synthesize heterogeneous user data through a psychometrics-based pipeline, aiming for realistic evaluation without using real private histories. In tests with 3 language models, 5 memory systems, and 10 synthetic users, systems did well on explicit semantic retrieval but declined on episodic tasks. Performance dropped further when the task required integrating fragmented information over time. ArXiv · AI/CL/LG's note

score 5

Categories: Research