On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
The paper finds that memory-based self-improving agents can look better or worse depending on run noise and task order.
The authors re-evaluate two such methods across multiple runs and shuffled task sequences. They report that complex, multi-step agent evaluations are already noisy, and self-improvement loops can amplify that variance. Default task orderings may act like hidden curricula, making gains depend on an implicit sequence. Adding rubrics and environment feedback to memory construction helps, but does not remove the remaining performance gaps. ArXiv · AI/CL/LG's note
The authors re-evaluate two such methods across multiple runs and shuffled task sequences. They report that complex, multi-step agent evaluations are already noisy, and self-improvement loops can amplify that variance. Default task orderings may act like hidden curricula, making gains depend on an implicit sequence. Adding rubrics and environment feedback to memory construction helps, but does not remove the remaining performance gaps. ArXiv · AI/CL/LG's note
score 5