AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization
AUSO trains agents to treat skills as action-level help, not permanent instructions.
The paper proposes a reinforcement learning method that starts by learning from teacher guidance and environment outcomes together, then shifts toward outcome-based policy optimization. In later training, it compares sampled actions with and without skill conditioning, strengthening actions where the skill helps and suppressing ones where it hurts. The authors report gains over competitive baselines on ALFWorld, WebShop, and SearchQA, including out-of-distribution generalization. ArXiv · AI/CL/LG's note
The paper proposes a reinforcement learning method that starts by learning from teacher guidance and environment outcomes together, then shifts toward outcome-based policy optimization. In later training, it compares sampled actions with and without skill conditioning, strengthening actions where the skill helps and suppressing ones where it hurts. The authors report gains over competitive baselines on ALFWorld, WebShop, and SearchQA, including out-of-distribution generalization. ArXiv · AI/CL/LG's note
score 4