Hierarchical Continuous Diffusion Language Models
HC-DLM uses a continuous latent state as the persistent generator, with tokens repeatedly read out and fed back during denoising.
The paper positions this as a fix for diffusion language models that either decode parallel tokens independently or keep continuous states too detached from valid token sequences. Its training objective is derived from a variational bound on token likelihood. The authors report gains over matched-size discrete and continuous diffusion baselines on Sudoku, Countdown, and LM1B. ArXiv · AI/CL/LG's note
The paper positions this as a fix for diffusion language models that either decode parallel tokens independently or keep continuous states too detached from valid token sequences. Its training objective is derived from a variational bound on token likelihood. The authors report gains over matched-size discrete and continuous diffusion baselines on Sudoku, Countdown, and LM1B. ArXiv · AI/CL/LG's note
score 5