Video2Skill: From Streaming Experience to Reusable Embodied Skills
The benchmark finds that current VLMs struggle most with knowing when a manipulation deserves a new reusable skill.
Video2Skill tests whether models can watch sequential robot and human activity videos, build a persistent skill library, and use it later. Across 19 open-source VLMs, many grouped repeated transformations at near-chance level, and larger models did not reliably do better. Fine-tuning improved grouping, but trained models tended to consolidate familiar skills instead of expanding the library for unseen transformations. HF Daily Papers' note
Video2Skill tests whether models can watch sequential robot and human activity videos, build a persistent skill library, and use it later. Across 19 open-source VLMs, many grouped repeated transformations at near-chance level, and larger models did not reliably do better. Fine-tuning improved grouping, but trained models tended to consolidate familiar skills instead of expanding the library for unseen transformations. HF Daily Papers' note
score 4