RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents
The benchmark argues many robot policies are passing tasks by reading the scene, not the instruction.
RoboFollow targets “low scene entropy,” where a scene effectively allows one task and language can become decorative. Its tests make each scene support multiple distinct task branches, then perturb layout and semantics across four levels. Across nine VLA and WAM policies, strong baseline performance did not reliably survive those harder instruction checks. The authors say stronger VLM backbones and several training mitigations still failed to close the gap. HF Daily Papers' note
RoboFollow targets “low scene entropy,” where a scene effectively allows one task and language can become decorative. Its tests make each scene support multiple distinct task branches, then perturb layout and semantics across four levels. Across nine VLA and WAM policies, strong baseline performance did not reliably survive those harder instruction checks. The authors say stronger VLM backbones and several training mitigations still failed to close the gap. HF Daily Papers' note
score 6