Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
The paper proposes a full-song generator that separates high-level audio-token planning from flow-matching audio rendering.
Its framework covers lyrics-to-song, instrumental generation, and cover-song generation. The system uses a semantic-aware tokenizer, a hierarchical autoregressive model called hybird-LM, FullDiT for fidelity in VAE latent space, and a melody module for preserving reference melodies in covers. The authors also test reward-based post-training methods including DPO, GRPO, OPD, and flow-based GRPO. They report competitive results on a multilingual automatic benchmark and the Artificial Analysis Music with Vocals leaderboard. ArXiv · AI/CL/LG's note
Its framework covers lyrics-to-song, instrumental generation, and cover-song generation. The system uses a semantic-aware tokenizer, a hierarchical autoregressive model called hybird-LM, FullDiT for fidelity in VAE latent space, and a melody module for preserving reference melodies in covers. The authors also test reward-based post-training methods including DPO, GRPO, OPD, and flow-based GRPO. They report competitive results on a multilingual automatic benchmark and the Artificial Analysis Music with Vocals leaderboard. ArXiv · AI/CL/LG's note
score 5