Megadose Built for builders and researchers.

Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models

· HF Daily Papers ·
AE-VLA is presented as the strongest dual-arm policy tested for recombining known arm skills into unseen task compositions.

The paper introduces ACG-Bench, a dual-arm robotics benchmark with 23 task-condition pairs across 8 task families. It tests whether policies can satisfy task goals plus order, timing, and physical milestone constraints when familiar per-arm skills are recomposed. Using the same pi_0.5 backbone and shared evaluation setup, the authors compare augmentation and architecture choices, then combine arm-token grouping, SkillLoRA, and arm-wise attention into AE-VLA. AE-VLA reports 21.53% generalization success in simulation and 39.00% mean success on physical SO101 robots across five unseen conditions, both above the listed baselines.

HF Daily Papers' note

score 5

Categories: Research