Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking
A small dose of poisoned agent memory cut benchmark accuracy from 0.850 to 0.300.
The paper says simple false assertions, added to 1.2% of a LongMemEval corpus, were enough to make persistent memory retrieve bad information later. A four-stage write-time screening system caught indirect prompt-injection cases but rejected none of the 360 poisoned memories. Provenance weighting also failed as a clean fix: weak weighting behaved like no defense, while strong weighting blocked legitimate untrusted evidence. The author argues for bounded retrieval occupancy instead of additive provenance penalties. ArXiv · AI/CL/LG's note
The paper says simple false assertions, added to 1.2% of a LongMemEval corpus, were enough to make persistent memory retrieve bad information later. A four-stage write-time screening system caught indirect prompt-injection cases but rejected none of the 360 poisoned memories. Provenance weighting also failed as a clean fix: weak weighting behaved like no defense, while strong weighting blocked legitimate untrusted evidence. The author argues for bounded retrieval occupancy instead of additive provenance penalties. ArXiv · AI/CL/LG's note
score 6