Megadose Built for builders and researchers.

Representation-Space MMD for Diffusion Language Models

· HF Daily Papers ·
The paper proposes a lighter post-training objective for diffusion language models by matching generated and reference features inside a frozen DLM.

It estimates MMD from contextual token-position features, getting multiple observations from one extractor pass. The method uses policy gradients for discrete models and direct latent differentiation for continuous ones. Reported results include lower generative perplexity on OpenWebText, better accuracy-computation trade-offs on GSM8K, and more parallel decoding on 16B DMax-LLaDA2.0 models with similar or higher math and code accuracy. HF Daily Papers' note

score 4

Categories: Research