What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies
The paper pins robot imitation failures on phase-specific visual grounding errors, not lost motor skill.
Using ACT, the authors add controlled distractor objects and receptacles and trace failures to picking and placement. They find sensitivity depends on both the distractor’s color or shape similarity and the manipulation stage. Distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting improve target selection while preserving control-relevant spatial information. The gains hold in simulation, on a physical UR3e, and in a pretrained vision-language-action policy for state-conditioned instrument handling. ArXiv · AI/CL/LG's note
Using ACT, the authors add controlled distractor objects and receptacles and trace failures to picking and placement. They find sensitivity depends on both the distractor’s color or shape similarity and the manipulation stage. Distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting improve target selection while preserving control-relevant spatial information. The gains hold in simulation, on a physical UR3e, and in a pretrained vision-language-action policy for state-conditioned instrument handling. ArXiv · AI/CL/LG's note
score 4