Transferring the Intelligence of VLMs to Robotic Control
RoboDawn lets an agentic VLM control a robot through discrete motion and gripper commands, without task-specific robot training.
The system closes the loop: the model observes the visual state, reasons about the next action, executes it, then adjusts from the result. A few in-context demonstrations teach both the interface and task strategy. On RoboTwin 2.0 C2R, success rose from 53.2% zero-shot to 73.6% one-shot, above the cited π0.5 baseline at 46.0%. The authors also report transfer to real Franka robots for block-in-basket and block-stacking tasks. HF Daily Papers' note
The system closes the loop: the model observes the visual state, reasons about the next action, executes it, then adjusts from the result. A few in-context demonstrations teach both the interface and task strategy. On RoboTwin 2.0 C2R, success rose from 53.2% zero-shot to 73.6% one-shot, above the cited π0.5 baseline at 46.0%. The authors also report transfer to real Franka robots for block-in-basket and block-stacking tasks. HF Daily Papers' note
score 5