PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
The benchmark tests whether personal agents actually get better from retained experience, not just whether they store it.
PAST-Bench runs agents through ordered fresh-session tasks with retained experience switched on and off under matched conditions. It covers 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and updates. The authors report that gains are real but uneven, and that similar headline improvements can come from different underlying save-retrieve-update behavior. They also introduce Hermes+, which improves average gains and gives clearer pathway evidence, especially when agents must replace outdated state. HF Daily Papers' note
PAST-Bench runs agents through ordered fresh-session tasks with retained experience switched on and off under matched conditions. It covers 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and updates. The authors report that gains are real but uneven, and that similar headline improvements can come from different underlying save-retrieve-update behavior. They also introduce Hermes+, which improves average gains and gives clearer pathway evidence, especially when agents must replace outdated state. HF Daily Papers' note
score 5