Megadose Built for builders and researchers.

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

· HF Daily Papers ·
The paper argues VLA robots can look competent while acting on shortcut cues instead of task evidence.

PerturBot targets those shortcuts during training with wrist-view perturbations, richer decision-relevant captions, and relabeled random or failed trajectory segments. The authors say more demonstrations can improve success rates without fixing the underlying reliance on visual, language, or motor priors. They also propose GroundingFscore as an offline check for whether a policy is using task evidence rather than those shortcuts. HF Daily Papers' note

score 4

Categories: Research