StepAudio 3 Music Technical Report
StepAudio’s new music model uses an explicit ABC-notation planning step before generating long-form audio.
The report describes a system that can generate songs, instrumentals, accompaniments from dry vocals, and cover-song synthesis up to 5 minutes and 30 seconds. Its tokenizer compresses music into a 50-Hz single-codebook stream, while a diffusion Transformer predicts continuous latents decoded into 48-kHz audio. The authors say reinforcement learning with DPO improved preference-based results, with top scores among tested systems on several AudioBox and MuQ-MuLan measures. On the preliminary Artificial Analysis Music Arena Vocals leaderboard, it ranks behind Suno V5.5 and Mureka, and ahead of Suno V5 and several other models. HF Daily Papers' note
The report describes a system that can generate songs, instrumentals, accompaniments from dry vocals, and cover-song synthesis up to 5 minutes and 30 seconds. Its tokenizer compresses music into a 50-Hz single-codebook stream, while a diffusion Transformer predicts continuous latents decoded into 48-kHz audio. The authors say reinforcement learning with DPO improved preference-based results, with top scores among tested systems on several AudioBox and MuQ-MuLan measures. On the preliminary Artificial Analysis Music Arena Vocals leaderboard, it ranks behind Suno V5.5 and Mureka, and ahead of Suno V5 and several other models. HF Daily Papers' note
score 6