SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
SkillProx reports a 3.0-point average accuracy gain over the strongest gradient-based skill-refinement baseline.
The paper frames agent “skills” as reusable text artifacts that can be improved without changing model weights. Its method pairs diagnosis-driven edits with rollback checks, then audits smaller knowledge units for whether they should be kept, demoted, consolidated, or removed. The authors say tests across in-distribution and out-of-distribution benchmarks, using multiple backbone LLMs, show the two stages contribute complementary gains. ArXiv · AI/CL/LG's note
The paper frames agent “skills” as reusable text artifacts that can be improved without changing model weights. Its method pairs diagnosis-driven edits with rollback checks, then audits smaller knowledge units for whether they should be kept, demoted, consolidated, or removed. The authors say tests across in-distribution and out-of-distribution benchmarks, using multiple backbone LLMs, show the two stages contribute complementary gains. ArXiv · AI/CL/LG's note
score 4