Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
The paper argues that memory usefulness cannot be measured if the system never retrieves the memory in the first place.
Behnam and Wang frame this as a retrieval-level positivity failure: store-level tests look identical when a memory is never surfaced. Their Causal Memory Policy reserves context slots for sampled memories with known propensities, then estimates utility with inverse propensity weighting. In their experiments, identification fails for 54% of required memories on LongMemEval and 67% on LoCoMo, while CMP raises required-vs-non-required discrimination from 0.54 to 0.66 AUC. They also find that query-specific utility does not reliably aggregate into a retention rule for unseen queries. ArXiv · AI/CL/LG's note
Behnam and Wang frame this as a retrieval-level positivity failure: store-level tests look identical when a memory is never surfaced. Their Causal Memory Policy reserves context slots for sampled memories with known propensities, then estimates utility with inverse propensity weighting. In their experiments, identification fails for 54% of required memories on LongMemEval and 67% on LoCoMo, while CMP raises required-vs-non-required discrimination from 0.54 to 0.66 AUC. They also find that query-specific utility does not reliably aggregate into a retention rule for unseen queries. ArXiv · AI/CL/LG's note
score 5