AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
AuK packages speech generation and editing into one open-source model interface driven by instructions and audio context.
The report says the team built about 3.03 billion instruction-audio instances and 1.95 million hours of effective supervision across five speech and audio task families. AuK uses a multimodal language model, a jointly trained audio VAE, and a hybrid rectified-flow Transformer for generation. A distilled version, AuK-Flash, runs 4-step inference without classifier-free guidance and is reported at 4.5x faster than the full model under matched conditions. The authors say they are releasing source code and model weights for reproducibility and further research. HF Daily Papers' note
The report says the team built about 3.03 billion instruction-audio instances and 1.95 million hours of effective supervision across five speech and audio task families. AuK uses a multimodal language model, a jointly trained audio VAE, and a hybrid rectified-flow Transformer for generation. A distilled version, AuK-Flash, runs 4-step inference without classifier-free guidance and is reported at 4.5x faster than the full model under matched conditions. The authors say they are releasing source code and model weights for reproducibility and further research. HF Daily Papers' note
score 6