Megadose AI progress, ranked and analyzed.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

· HF Daily Papers ·
SwanTale is built to generate reusable multi-speaker voices, scenes, and effects from either text instructions or reference audio.

ByteDance’s technical report pairs a new captioned dataset, SwanData-Caption, with a model meant for expressive speech and broader audio generation. The instruct mode uses captions for environment, speaker style, and fine-grained content; the zero-shot mode adds reference audio. The authors say SwanTale leads on several zero-shot and instruct metrics, including expressiveness, and can handle complex multi-speaker audio instructions. HF Daily Papers' note

score 5

Categories: Research