Megadose Built for builders and researchers.

ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills

· HF Daily Papers ·
ViSkill turns successful visual-agent runs into reusable visual skill cards, then feeds them back into training.

The paper argues that text-based skill memories lose spatial structure that VLM agents need. ViSkill stores successful interactions as composite visual cards, retrieves them during inference, and also uses them for reward shaping. New successful trajectories are distilled back into the skill library, creating a loop between policy improvement and skill accumulation. In tests on Sokoban, FrozenLake, and PrimitiveSkill, it reports a 0.89 overall success rate, or 0.91 with cold-start initialization, ahead of the evaluated baselines.

HF Daily Papers' note

score 4

Categories: Research