Megadose Built for builders and researchers.

Enhancing Diffusion Language Models with Autoregressive Post-Training Weights

· ArXiv · AI/CL/LG ·
AR post-training updates can be reused directly to improve diffusion language models.

The paper says adding existing autoregressive post-training weight updates to diffusion base models brings performance close to direct diffusion post-training. It proposes A2D, a training-free method that composes AR and diffusion post-training updates. The authors report gains across several dLLMs on instruction following, math reasoning, and coding, without extra training or inference-time compute. ArXiv · AI/CL/LG's note

score 5

Categories: Research