Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
Recuris reports broad gains by making agent memory update itself through validated failure evidence.
The paper introduces a working-memory layer that tracks task state and selects skills from experiential memory without relying on the full interaction history. A fixed meta-agent uses execution evidence to make localized, validation-gated updates to skill memory. Across four long-horizon benchmarks and ten models, it improved 35 of 37 completed model-benchmark pairs. The authors report larger gains on longer tasks and reductions in common long-horizon failures of up to 80%. ArXiv · AI/CL/LG's note
The paper introduces a working-memory layer that tracks task state and selects skills from experiential memory without relying on the full interaction history. A fixed meta-agent uses execution evidence to make localized, validation-gated updates to skill memory. Across four long-horizon benchmarks and ten models, it improved 35 of 37 completed model-benchmark pairs. The authors report larger gains on longer tasks and reductions in common long-horizon failures of up to 80%. ArXiv · AI/CL/LG's note
score 5