VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
The paper claims VLA-Precision gets real robots to 98.3% mean success on precision chemistry tasks after under an hour of online RL per task.
The framework targets pretrained vision-language-action models that struggle with repeatable, high-precision manipulation. Its ACoB method mixes early intervention-guided learning with later return propagation and preference ranking to improve policy behavior while limiting drift. The authors also introduce ACoB-Stream to cut large-model overhead, reporting up to 10.9x gains in throughput and computational efficiency. Tests span nine chemistry tasks, four categories, and four robot embodiments, with 27.6-second episodes running faster than the VLA and RL baselines. HF Daily Papers' note
The framework targets pretrained vision-language-action models that struggle with repeatable, high-precision manipulation. Its ACoB method mixes early intervention-guided learning with later return propagation and preference ranking to improve policy behavior while limiting drift. The authors also introduce ACoB-Stream to cut large-model overhead, reporting up to 10.9x gains in throughput and computational efficiency. Tests span nine chemistry tasks, four categories, and four robot embodiments, with 27.6-second episodes running faster than the VLA and RL baselines. HF Daily Papers' note
score 5