SPyCE: Skill-Policy Co-evolution for Multimodal Agents
The paper proposes training multimodal agents to turn useful reasoning trajectories into reusable tool-use skills, then feed those skills back into policy learning.
SPyCE builds a hierarchical skill library: execution skills for local visual operations and workflow skills for higher-level tool orchestration. During reinforcement learning, the policy conditions on retrieved skills, while stronger policy rollouts update the library. The authors report gains over RL-based and memory-based baselines across eight benchmarks, with analysis pointing to both the hierarchy and the co-evolution loop as necessary pieces. ArXiv · AI/CL/LG's note
SPyCE builds a hierarchical skill library: execution skills for local visual operations and workflow skills for higher-level tool orchestration. During reinforcement learning, the policy conditions on retrieved skills, while stronger policy rollouts update the library. The authors report gains over RL-based and memory-based baselines across eight benchmarks, with analysis pointing to both the hierarchy and the co-evolution loop as necessary pieces. ArXiv · AI/CL/LG's note
score 5