Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
GPD trains an RGB-only VLM with privileged 3D cues, then removes those cues at deployment.
The paper targets spatial failures that come from bad depth or direction judgments in RGB inputs. Its method renders depth, semantic, and bird’s-eye-view cues as compact text for a teacher during on-policy self-distillation. The extra distillation loss is applied only on incorrect trajectories, alongside GRPO. On a 4B backbone, the authors report gains over GRPO and answer-privileged OPSD across several spatial reasoning benchmarks. ArXiv · AI/CL/LG's note
The paper targets spatial failures that come from bad depth or direction judgments in RGB inputs. Its method renders depth, semantic, and bird’s-eye-view cues as compact text for a teacher during on-policy self-distillation. The extra distillation loss is applied only on incorrect trajectories, alongside GRPO. On a 4B backbone, the authors report gains over GRPO and answer-privileged OPSD across several spatial reasoning benchmarks. ArXiv · AI/CL/LG's note
score 5