SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
SKILLER trains small-model agents to generate reusable task skills through natural-language reinforcement signals.
The paper uses a stronger model as both actor and critic while the smaller agent system serves as the environment. Tested on Qwen3.5-9B and Qwen3.5-4B across five benchmarks, SKILLER beat four skill-generation or evolution baselines. Reported gains ranged from 4.3 to 20.4 percentage points for the 9B model and 1.8 to 13.3 points for the 4B model. On single-skill tasks in SkillsBench, it matched strong closed-source model performance. HF Daily Papers' note
The paper uses a stronger model as both actor and critic while the smaller agent system serves as the environment. Tested on Qwen3.5-9B and Qwen3.5-4B across five benchmarks, SKILLER beat four skill-generation or evolution baselines. Reported gains ranged from 4.3 to 20.4 percentage points for the 9B model and 1.8 to 13.3 points for the 4B model. On single-skill tasks in SkillsBench, it matched strong closed-source model performance. HF Daily Papers' note
score 4