Megadose AI progress, ranked and analyzed.

OPD-V: Visual On-Policy Self-Distillation with Modality Balance

· ArXiv · AI/CL/LG ·
OPD-V treats modality balance itself as the signal for improving visual reasoning distillation.

The paper argues that visual self-distillation in multimodal models is weakened when text dominates generation and image information is underused. It tests this by comparing a zoom-in “Positive Teacher” and a masked-image “Negative Teacher,” then reading the effect through reasoning correctness and token logits. OPD-V uses those teachers to define a modality-balance trust region for selecting on-policy tokens. Across 6 benchmarks, 4 MLLM backbones, and 5 post-training methods, the authors report consistent reasoning gains with lower training cost. ArXiv · AI/CL/LG's note

score 4

Categories: Research