Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Pivot-SD trains masked diffusion models on the few denoising commitments that most reduce uncertainty.
The paper frames those commitments as “pivots,” selected by an information-gain metric over the still-masked positions. Successful pivots are reinforced with cross-entropy; failed pivots are handled with targeted unlikelihood while the rest of the failed trajectory is left alone. With 200 questions and four rollouts each, the authors report gains for LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines on math and code benchmarks. HF Daily Papers' note
The paper frames those commitments as “pivots,” selected by an information-gain metric over the still-masked positions. Successful pivots are reinforced with cross-entropy; failed pivots are handled with targeted unlikelihood while the rest of the failed trajectory is left alone. With 200 questions and four rollouts each, the authors report gains for LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines on math and code benchmarks. HF Daily Papers' note
score 5