SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
SkillGate trains the skill-picking tokens separately from the rest of the agent run.
The paper argues that ordinary outcome-reward RL misassigns credit in long-horizon agents: the chosen-skill tokens get little signal, and that signal can turn negative when later execution fails. Its proposed fix splits credit into two channels, with outcome credit for execution and action-local credit for the skill read. On five agentic benchmarks with a 16-skill slate, the authors report trial success rising from 40.8% to 53.2% for a 9B policy. HF Daily Papers' note
The paper argues that ordinary outcome-reward RL misassigns credit in long-horizon agents: the chosen-skill tokens get little signal, and that signal can turn negative when later execution fails. Its proposed fix splits credit into two channels, with outcome credit for execution and action-local credit for the skill read. On five agentic benchmarks with a 16-skill slate, the authors report trial success rising from 40.8% to 53.2% for a 9B policy. HF Daily Papers' note
score 5