Megadose AI progress, ranked and analyzed.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

· HF Daily Papers ·
VAD tries to separate teacher corrections backed by visual evidence from corrections driven by language priors.

The method compares a fixed teacher with the relevant visual evidence present and removed, then uses the probability shift to estimate which token corrections the evidence supports or refutes. It reconstructs the student’s training target from that visually attributable component, with the privileged teacher kept as a weak regularizer. The paper reports gains over direct privileged-view distillation and visual-advantage weighting across six fine-grained visual benchmarks at 4B and 9B scales. HF Daily Papers' note

score 4

Categories: Research