Qwen-Music Technical Report
Qwen-Music is presented as a full-song generator that plans melodies before rendering high-fidelity stereo audio.
The report describes a three-part system: tokenizer, LLM, and renderer. Its Melody-CoT mechanism plans melody tokens before generating the song, which the authors say improves creativity, structure, and melody cloning from reference audio. The model was trained on more than 5 million hours of multilingual music, then post-trained with supervised initialization, offline DPO, and online GSPO. In tests across 600 Chinese and English prompts, the authors report state-of-the-art results on 13 of 16 objective metrics and professional-evaluator preference over leading proprietary systems. HF Daily Papers' note
The report describes a three-part system: tokenizer, LLM, and renderer. Its Melody-CoT mechanism plans melody tokens before generating the song, which the authors say improves creativity, structure, and melody cloning from reference audio. The model was trained on more than 5 million hours of multilingual music, then post-trained with supervised initialization, offline DPO, and online GSPO. In tests across 600 Chinese and English prompts, the authors report state-of-the-art results on 13 of 16 objective metrics and professional-evaluator preference over leading proprietary systems. HF Daily Papers' note
score 7