Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
The paper claims synthetic spatial guidance can teach MLLMs perceptual skills that transfer to real images.
Where-OPD trains a teacher model with privileged textual cues about relevant objects and coordinates in procedurally generated scenes, while the student must answer from only the image and question. The authors say this improves counting, document, and chart understanding across multiple models. They report a 3.23-point average gain on real-world perception benchmarks despite post-training only on synthetic scenes. ArXiv · AI/CL/LG's note
Where-OPD trains a teacher model with privileged textual cues about relevant objects and coordinates in procedurally generated scenes, while the student must answer from only the image and question. The authors say this improves counting, document, and chart understanding across multiple models. They report a 3.23-point average gain on real-world perception benchmarks despite post-training only on synthetic scenes. ArXiv · AI/CL/LG's note
score 4