Megadose Built for builders and researchers.

Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models

· ArXiv · AI/CL/LG ·
A robot policy that succeeds on one wording can fail almost completely on a near-synonym.

The paper reports that VLA models are sharply affected by small instruction edits, including a LIBERO stove task where success drops from 100% for “switch on the stove” to 2% for “switch on the hot plate.” The authors test single-edit swings and phrase search, then mitigate the issue by using an LLM to turn evidence from training-task phrasings into rephrasing rules. At deployment, each instruction is rewritten once under those rules, without retraining the policy or adding per-step checks. The frozen policy improves 16% to 27% relative on twelve held-out tasks, with most gains on out-of-distribution tasks. ArXiv · AI/CL/LG's note

score 5

Categories: Research