FastOPD: On-Policy Distillation for Lightweight VLA Deployment
FastOPD distills large VLA robot policies into compact students that can run with far fewer inference steps.
The paper says its on-policy distillation method trains a lightweight student to learn teacher dynamics using single-state supervision and a self-consistency objective. On LIBERO, it retained 84% of π₀.₅ performance with two inference steps while cutting latency by 78.1%. With LingBot-VLA as teacher, it improved single-step success over the base student by 15.9 points on RoboTwin 2.0. The authors also report tests with a World Action Model and a real-robot student distilled from MolmoAct2. HF Daily Papers' note
The paper says its on-policy distillation method trains a lightweight student to learn teacher dynamics using single-state supervision and a self-consistency objective. On LIBERO, it retained 84% of π₀.₅ performance with two inference steps while cutting latency by 78.1%. With LingBot-VLA as teacher, it improved single-step success over the base student by 15.9 points on RoboTwin 2.0. The authors also report tests with a World Action Model and a real-robot student distilled from MolmoAct2. HF Daily Papers' note
score 4