Agent Priors-guided Policy Learning
APPL makes a robot skill’s training assumptions part of the interface used to compose that skill later.
The paper argues that composition loses useful information when skills are exposed only as names, instructions, or symbolic operators. APPL has a construction agent split demonstrations into reusable skills, propose structural priors for each one, then train and verify a policy for each prior. At runtime, another agent chooses among those prior-specific policies and composes them toward new goals. The authors report gains on MetaWorld and long-horizon ManiSkill tasks, including better out-of-distribution skill generalization and unseen skill compositions. HF Daily Papers' note
The paper argues that composition loses useful information when skills are exposed only as names, instructions, or symbolic operators. APPL has a construction agent split demonstrations into reusable skills, propose structural priors for each one, then train and verify a policy for each prior. At runtime, another agent chooses among those prior-specific policies and composes them toward new goals. The authors report gains on MetaWorld and long-horizon ManiSkill tasks, including better out-of-distribution skill generalization and unseen skill compositions. HF Daily Papers' note
score 4