SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models
The paper argues pruned visual tokens still contain usable evidence, but models fail to use it reliably.
SCOPD trains a sparse-token student against a full-context teacher on the student’s own reasoning prefixes, without ground-truth answers or inference-time changes. The authors call the failure mode a “representation-utilization gap,” shown by Pass@K gains from repeated sampling on the same pruned representation. At 10% visual-token retention, SCOPD improves retained unpruned performance from 86.37% to 90.49%, and SCOPD+ raises it to 92.43% across 13 benchmarks. HF Daily Papers' note
SCOPD trains a sparse-token student against a full-context teacher on the student’s own reasoning prefixes, without ground-truth answers or inference-time changes. The authors call the failure mode a “representation-utilization gap,” shown by Pass@K gains from repeated sampling on the same pruned representation. At 10% visual-token retention, SCOPD improves retained unpruned performance from 86.37% to 90.49%, and SCOPD+ raises it to 92.43% across 13 benchmarks. HF Daily Papers' note
score 5