Megadose AI progress, ranked and analyzed.

On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data

· ArXiv · AI/CL/LG ·
OnPoKD makes distillation targets a learned, sample-by-sample policy decision instead of a fixed teacher output.

The paper says current vision-language distillation can become unreliable when classes or domains shift. Its framework uses cues from the teacher, student, and zero-shot prior to balance teacher supervision, prior guidance, and hard labels during training. The controller is trained with validation feedback and is removed at inference, so the student keeps the original test-time cost. The authors report gains over strong baselines on base-to-novel generalization and cross-dataset transfer benchmarks. ArXiv · AI/CL/LG's note

score 5

Categories: Research