SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
The paper’s core claim is that models need verified practice using skills, not just access to them.
SKT builds synthetic tasks and executable trajectories from 2,000 public agent skills, keeping only runs that successfully use every required skill. The resulting set includes 4,000 task packages and 27,164 verified trajectories. The authors also introduce SkillEval, a held-out benchmark made from a separate test pool. Across their experiments, supervised fine-tuning on SKT trajectories improves skill-use performance, with gains tied to verification quality and broader skill coverage. HF Daily Papers' note
SKT builds synthetic tasks and executable trajectories from 2,000 public agent skills, keeping only runs that successfully use every required skill. The resulting set includes 4,000 task packages and 27,164 verified trajectories. The authors also introduce SkillEval, a held-out benchmark made from a separate test pool. Across their experiments, supervised fine-tuning on SKT trajectories improves skill-use performance, with gains tied to verification quality and broader skill coverage. HF Daily Papers' note
score 5