Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
The paper proposes “skills” as the verification layer for open-ended LLM self-training.
Skill Self-Play uses a proposer, solver, and dynamic skill controller in a reinforcement-learning loop. The proposer generates tasks around sampled skills, the solver attempts them, and the controller uses execution feedback to update the skill library. The authors say this preserves verifiable feedback without locking training into narrow environments. Evaluations on tool-use and reasoning benchmarks are reported as improving strong models and reversing failures in initially misaligned ones. HF Daily Papers' note
Skill Self-Play uses a proposer, solver, and dynamic skill controller in a reinforcement-learning loop. The proposer generates tasks around sampled skills, the solver attempts them, and the controller uses execution feedback to update the skill library. The authors say this preserves verifiable feedback without locking training into narrow environments. Evaluations on tool-use and reasoning benchmarks are reported as improving strong models and reversing failures in initially misaligned ones. HF Daily Papers' note
score 5