Megadose AI progress, ranked and analyzed.

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

· ArXiv · AI/CL/LG ·
SAGE trains a small RL policy to use VLM help during learning, then run without it.

The framework queries the VLM only when the learner is uncertain, executes the suggested action in training, and distills that guidance into the policy. It can downweight bad teacher actions using environment-derived advantages instead of trusting every suggestion equally. In sparse-reward visual reasoning and navigation tasks, it improved over unguided RL in several environments and sometimes exceeded the VLM teacher. It also cuts VLM use by limiting calls during training and requiring none at deployment. ArXiv · AI/CL/LG's note

score 5

Categories: Research