Alpha Diffusion Language Models: Factorization Alone Is Not the Problem
AlphaDLM changes the training loss so parallel token generation can keep valid joint predictions with fewer denoising steps.
The paper argues that cross-entropy fits token marginals, while parallel diffusion decoding needs consistency across tokens. Its sequence-level alpha loss moves between cross-entropy behavior and a joint-mode objective. On TinyGSM, the method reports 34.6% GSM8K accuracy using four model evaluations, with additional tests on SDAR-1.7B for code and math. ArXiv · AI/CL/LG's note
The paper argues that cross-entropy fits token marginals, while parallel diffusion decoding needs consistency across tokens. Its sequence-level alpha loss moves between cross-entropy behavior and a joint-mode objective. On TinyGSM, the method reports 34.6% GSM8K accuracy using four model evaluations, with additional tests on SDAR-1.7B for code and math. ArXiv · AI/CL/LG's note
score 5