Megadose AI progress, ranked and analyzed.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

· ArXiv · AI/CL/LG ·
VAD tries to separate teacher corrections backed by visual evidence from corrections driven by language priors or teacher quirks.

The paper introduces Visual Attribution Distillation, which compares a fixed teacher’s token probabilities with relevant visual evidence present and removed. That counterfactual shift is used as a proxy for whether the evidence supports or refutes candidate tokens. The method projects the teacher correction onto that visual-evidence direction, then uses the reconstructed target as the main training signal. The authors report gains over direct privileged-view distillation and visual-advantage weighting across six fine-grained visual benchmarks at 4B and 9B scale. ArXiv · AI/CL/LG's note

score 4

Categories: Research