dQwen3.5: Hybrid-Attention Diffusion Language Models
Hybrid Qwen3.5 backbones adapted into diffusion language models reached the same training loss as a full-attention control using about half the tokens.
The paper adapts Qwen3.5 at 0.8B, 2B, 4B, and 9B scales into a dQwen3.5 family. Its central problem is whether AR hybrid architectures that mix attention and RNN layers can be made bidirectional enough for diffusion language modeling. The authors report that the resulting models behave like full-attention DLMs in any-order decoding and perform strongly with parallel decoding. ArXiv · AI/CL/LG's note
The paper adapts Qwen3.5 at 0.8B, 2B, 4B, and 9B scales into a dQwen3.5 family. Its central problem is whether AR hybrid architectures that mix attention and RNN layers can be made bidirectional enough for diffusion language modeling. The authors report that the resulting models behave like full-attention DLMs in any-order decoding and perform strongly with parallel decoding. ArXiv · AI/CL/LG's note
score 5