Megadose AI progress, ranked and analyzed.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

· HF Daily Papers ·
Anchor-Align is meant to keep robot-policy finetuning from erasing the pretrained VLM features that help generalization.

The paper says standard behavior-cloning finetuning can overwrite visual and semantic representations in VLA policies. Its method adds layer-wise distillation from a frozen VLM and trains language and action prediction on the same robot observation using discrete motion-direction labels. On a physical xArm7 setup, it reports success-rate gains from 28% to 54% and from 37% to 60% across two VLA architectures. The authors also report simulation gains on OOD perturbations, perceptual robustness, and long-horizon control.

HF Daily Papers' note

score 4

Categories: Research