Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation
The paper targets a specific OPD failure: students get better at pass@1 but stop gaining the teacher’s sampling diversity.
The authors trace that collapse to entropy-contracting token updates, using a first-order influence proxy to identify where they happen. Their IDA-OPD method keeps entropy-expanding updates and shrinks the harmful ones without requiring full-vocabulary teacher probabilities. In reasoning distillation experiments, they report better pass@$k$ while largely preserving vanilla OPD’s pass@1. Source: HF Daily Papers' note.
The authors trace that collapse to entropy-contracting token updates, using a first-order influence proxy to identify where they happen. Their IDA-OPD method keeps entropy-expanding updates and shrinks the harmful ones without requiring full-vocabulary teacher probabilities. In reasoning distillation experiments, they report better pass@$k$ while largely preserving vanilla OPD’s pass@1. Source: HF Daily Papers' note.
score 5