Megadose AI progress, ranked and analyzed.

Distribution Matching Distillation for Continuous Diffusion Language Models

· ArXiv · AI/CL/LG ·
The paper reports lower perplexity for continuous diffusion language models by distilling them into cheaper multi-step generators.

The authors compare two reverse-KL distillation methods using the same student architecture: Simplex-DMD for continuous token relaxations and Reinforce-DMD for categorical sampling. On OpenWebText at 1,024-token length, Simplex-DMD reaches 45.6 generative perplexity in 4 NFEs, which the paper says is 49% below the strongest evaluated diffusion baseline at matched entropy and budget. Reinforce-DMD does better at larger budgets, reaching 14.9 perplexity with 256 NFEs, a reported 20% reduction under the same protocol. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research