Megadose AI progress, ranked and analyzed.

Self-Supervised Visual On-Policy Distillation

· HF Daily Papers ·
The paper’s trick is to create teacher-student asymmetry by weakening the student’s view, not by giving the teacher privileged information.

S²VOPD trains on the teacher’s distribution from the original image while the student sees a strongly augmented version of the same image. The authors report that asymmetric augmentation helps, symmetric self-distillation hurts, and overly destructive augmentation can erase the evidence needed for the task. Across six fine-grained perception benchmarks, they say Qwen3.5-4B rises from 70.7% to 77.4% with the same training data. HF Daily Papers' note

score 4

Categories: Research