When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
Raw-turn selection matched extraction under tight memory budgets, with far lower write cost.
The authors tested agent memory on held-out LoCoMo conversations and LongMemEval. In their pre-registered setup, Jev-selected raw turns were non-inferior to an LLM-extracted memory on tight LoCoMo budgets, while costing 3,061 times less to write. Reranking helped most when only three of 30 candidates could be kept, but its gain nearly vanished at larger budgets, where extraction systems were more accurate. At matched context, Jev selected about as accurately as an LLM reranker with lower latency, though reranking reduced correct abstention. HF Daily Papers' note
The authors tested agent memory on held-out LoCoMo conversations and LongMemEval. In their pre-registered setup, Jev-selected raw turns were non-inferior to an LLM-extracted memory on tight LoCoMo budgets, while costing 3,061 times less to write. Reranking helped most when only three of 30 candidates could be kept, but its gain nearly vanished at larger budgets, where extraction systems were more accurate. At matched context, Jev selected about as accurately as an LLM reranker with lower latency, though reranking reduced correct abstention. HF Daily Papers' note
score 4