Megadose Built for builders and researchers.

ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills

· ArXiv · AI/CL/LG ·
ViSkill turns successful VLM agent runs into reusable visual skill cards, then feeds them back into training.

The paper says text-based skill methods lose spatial structure by converting layouts and actions into language. ViSkill stores successful interactions as visual-native cards that agents can retrieve for inference and reward shaping. New successful trajectories are distilled back into the skill library, creating a closed loop between skill growth and policy improvement. On Sokoban, FrozenLake, and PrimitiveSkill, the authors report a 0.89 overall success rate, or 0.91 with cold-start initialization. ArXiv · AI/CL/LG's note

score 5

Categories: Research