Megadose Built for builders and researchers.

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

· HF Daily Papers ·
E-MoE uses MoE routing decisions as a shared discrete latent to improve few-step masked diffusion generation.

The paper targets masked diffusion language models whose usual reverse process treats positions independently.
Its claim is that this factorization hurts sample quality in the low-step setting where diffusion decoding is supposed to be fastest.
E-MoE builds a mixture of factorized distributions using the routing choices from an MoE backbone, without adding active parameters versus the factorized baseline.
The authors report gains over factorized baselines on synthetic multimodal tests, binarized MNIST, and LM1B.
HF Daily Papers' note

score 4

Categories: Research