DolphinBench: Mapping the Pareto Frontier of Agent Memory
DolphinBench tests agent memory by whether it helps finish tasks, not by whether a model can answer prompted recall questions.
The benchmark uses three knowledge-work personas, each with about 500k tokens of user-message history. For each persona, it includes 200 tasks that depend on information from that history, verified by requiring an agent to succeed with the relevant context and fail without it. Submissions must report cost and latency alongside accuracy, so memory systems can be compared on tradeoffs as well as scores. Dataset and evaluation code are released with the paper. ArXiv · AI/CL/LG's note
The benchmark uses three knowledge-work personas, each with about 500k tokens of user-message history. For each persona, it includes 200 tasks that depend on information from that history, verified by requiring an agent to succeed with the relevant context and fail without it. Submissions must report cost and latency alongside accuracy, so memory systems can be compared on tradeoffs as well as scores. Dataset and evaluation code are released with the paper. ArXiv · AI/CL/LG's note
score 5