Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
For cross-tokenizer OPD, tighter supervision beat broader alignment in the paper’s tests.
The authors find that strict 1:1 token alignments already cover most student-generated tokens across three teacher-student pairs for math reasoning and code generation. A top-16 shared-vocabulary reverse-KL loss at those strict positions matched full shared-vocabulary OPD and beat the evaluated cross-tokenizer baselines. Adding span-level MSE supervision over mismatched groups filled the coverage gap but lowered accuracy, with gradient diagnostics suggesting weak or conflicting training signals. Source: HF Daily Papers' note.
The authors find that strict 1:1 token alignments already cover most student-generated tokens across three teacher-student pairs for math reasoning and code generation. A top-16 shared-vocabulary reverse-KL loss at those strict positions matched full shared-vocabulary OPD and beat the evaluated cross-tokenizer baselines. Adding span-level MSE supervision over mismatched groups filled the coverage gap but lowered accuracy, with gradient diagnostics suggesting weak or conflicting training signals. Source: HF Daily Papers' note.
score 4