Megadose Built for builders and researchers.

EXIMO: VLM Guided Exploration of VLA Policies

· HF Daily Papers ·
EXIMO uses a vision-language model as a planner so a VLA robot policy can gather better training data before fine-tuning.

The paper frames new-task learning for VLA manipulation policies as limited by costly teleoperation data and sample-inefficient RL. EXIMO splits training into explore, imitate, and optimize stages: planned exploration, supervised fine-tuning on collected task data, then residual off-policy RL. The authors say ablations show all three stages matter, with stronger sample efficiency and final performance than existing approaches. HF Daily Papers' note

score 4

Categories: Research