Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Where-OPD trains MLLMs with synthetic spatial hints, then removes the hints at test time.
The method gives a teacher model textual guidance about which objects and coordinates matter for a query, using procedurally generated scenes where that information is automatic. The student learns to match the teacher’s behavior from only the image and question. The paper reports gains on counting, document, and chart understanding benchmarks across multiple models. It also reports a 3.23-point average gain on six real-world perception benchmarks after post-training only on synthetic scenes. HF Daily Papers' note
The method gives a teacher model textual guidance about which objects and coordinates matter for a query, using procedurally generated scenes where that information is automatic. The student learns to match the teacher’s behavior from only the image and question. The paper reports gains on counting, document, and chart understanding benchmarks across multiple models. It also reports a 3.23-point average gain on six real-world perception benchmarks after post-training only on synthetic scenes. HF Daily Papers' note
score 4