StepAudio 3 Gen Technical Report
StepAudio 3 Gen is pitched as one model for TTS, voice design, vocals, sound effects, music, and mixed audio generation.
The report says the system uses a discrete autoregressive generator over RVQ audio tokens, rather than the diffusion Transformer approach common in recent general audio models. Its tokenizer represents general audio at 12.5 Hz in a shared 16-codebook space, with semantic and acoustic information preserved across layers. The authors identify progressive pretraining, an RVQ adaptor, and shared discrete autoregressive modeling as key design choices. They claim state-of-the-art results on TTS and voice design, while keeping broader speech, vocal, effects, and music generation ability. HF Daily Papers' note
The report says the system uses a discrete autoregressive generator over RVQ audio tokens, rather than the diffusion Transformer approach common in recent general audio models. Its tokenizer represents general audio at 12.5 Hz in a shared 16-codebook space, with semantic and acoustic information preserved across layers. The authors identify progressive pretraining, an RVQ adaptor, and shared discrete autoregressive modeling as key design choices. They claim state-of-the-art results on TTS and voice design, while keeping broader speech, vocal, effects, and music generation ability. HF Daily Papers' note
score 6