Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents
The paper finds that memory choice depends on the agent’s workload, with no substrate winning across the board.
The authors test memory-augmented agents across dense and sparse retrieval, text records, structural and hierarchical stores, refinement memories, parametric updates, and context-based mechanisms. Across three backbone models and four benchmark suites, broad retrieval helps long-context factual QA but can hurt sequential decision-making by pulling attention away from action-relevant context. They argue that scalable agent memory will need routing between substrates as history length and task demands change. HF Daily Papers' note
The authors test memory-augmented agents across dense and sparse retrieval, text records, structural and hierarchical stores, refinement memories, parametric updates, and context-based mechanisms. Across three backbone models and four benchmark suites, broad retrieval helps long-context factual QA but can hurt sequential decision-making by pulling attention away from action-relevant context. They argue that scalable agent memory will need routing between substrates as history length and task demands change. HF Daily Papers' note
score 5