ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
Sequential agents improved with practice, but often from context and feedback rather than durable new skills.
ContinualSkillBench tests in-context continual skill learning across five domains, with 100 linked subtasks in each. The paper reports that running tasks sequentially generally raises performance, though the size of the gain depends heavily on the model and domain. Explicit skill libraries did not beat plain in-context learning on average, but helped on tasks needing reusable procedures or precise outputs. Less capable models tended to build larger, more fragmented sets of task-specific skills. ArXiv · AI/CL/LG's note
ContinualSkillBench tests in-context continual skill learning across five domains, with 100 linked subtasks in each. The paper reports that running tasks sequentially generally raises performance, though the size of the gain depends heavily on the model and domain. Explicit skill libraries did not beat plain in-context learning on average, but helped on tasks needing reusable procedures or precise outputs. Less capable models tended to build larger, more fragmented sets of task-specific skills. ArXiv · AI/CL/LG's note
score 5