MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair
In the benchmark, poisoned long-term memories persisted in 84.2% of cases and drove the full attack chain in 50.3%.
MemSecBench tests agent memory systems through a controlled Write-Execute-Forget protocol across 310 cases in code, science, daily life, and office work. The authors compare 24 configurations spanning two agent harnesses, four memory backends, and three LLM backends. Among cases where poisoning succeeded, 59.6% completed the Execute chain, while selective repair succeeded in 56.1%. The paper says memory stacks differed sharply, including gaps of 16.1 percentage points in end-to-end attack success and 41.3 points in selective repair. ArXiv · AI/CL/LG's note
MemSecBench tests agent memory systems through a controlled Write-Execute-Forget protocol across 310 cases in code, science, daily life, and office work. The authors compare 24 configurations spanning two agent harnesses, four memory backends, and three LLM backends. Among cases where poisoning succeeded, 59.6% completed the Execute chain, while selective repair succeeded in 56.1%. The paper says memory stacks differed sharply, including gaps of 16.1 percentage points in end-to-end attack success and 41.3 points in selective repair. ArXiv · AI/CL/LG's note
score 5