EXIMO: VLM Guided Exploration of VLA Policies
EXIMO uses a vision-language model as a planner so a VLA robot policy can gather better training data before fine-tuning.
The paper frames new-task learning for VLA manipulation policies as limited by costly teleoperation data and sample-inefficient RL. EXIMO splits training into explore, imitate, and optimize stages: planned exploration, supervised fine-tuning on collected task data, then residual off-policy RL. The authors say ablations show all three stages matter, with stronger sample efficiency and final performance than existing approaches. HF Daily Papers' note
The paper frames new-task learning for VLA manipulation policies as limited by costly teleoperation data and sample-inefficient RL. EXIMO splits training into explore, imitate, and optimize stages: planned exploration, supervised fine-tuning on collected task data, then residual off-policy RL. The authors say ablations show all three stages matter, with stronger sample efficiency and final performance than existing approaches. HF Daily Papers' note
score 4