Megadose AI progress, ranked daily.

MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization

· HF Daily Papers ·
MA-VLA assigns explicit mid-level action prompts to each robot arm so unseen collaboration patterns can be recomposed at test time.

The paper says current VLA models often treat language as one global instruction, which makes arm-specific coordination hard to transfer. MA-VLA breaks cooperative behavior into atomic prompts and allocates them per arm, with “Arm Shuffle” training to reduce fixed role dependence. The authors report that prior state-of-the-art VLAs largely fail on collaboration patterns held out from training, while MA-VLA succeeds in simulation and real-world tests. HF Daily Papers' note

score 5

Categories: Research