Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Standard distillation appears to trade away factual recall during mid-training, even as reasoning keeps improving.
The paper reports controlled experiments showing forward KL distillation helps both reasoning and recall in pre-training, but slows factual recall acquisition during mid-training. The authors attribute that split to teacher confidence patterns and the student’s changing knowledge state. Their proposed Switch Distillation routes confident teacher predictions through distillation and falls back to cross-entropy otherwise. They report stronger reasoning and knowledge/commonsense results while preserving most factual recall, with the gains carrying through post-training. ArXiv · AI/CL/LG's note
The paper reports controlled experiments showing forward KL distillation helps both reasoning and recall in pre-training, but slows factual recall acquisition during mid-training. The authors attribute that split to teacher confidence patterns and the student’s changing knowledge state. Their proposed Switch Distillation routes confident teacher predictions through distillation and falls back to cross-entropy otherwise. They report stronger reasoning and knowledge/commonsense results while preserving most factual recall, with the gains carrying through post-training. ArXiv · AI/CL/LG's note
score 5