Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Tacit-TTS claims transcript-free voice cloning with more than 10x faster generation on longer utterances than IndexTTS2.
The paper replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation. It adds training-free acoustic length estimation and uses ReFlow distillation to speed the flow-matching renderer. The authors report competitive zero-shot quality across two English and two Mandarin datasets. They also test references in eight other languages, infant babble, and synthetic gibberish, arguing transcript-dependent systems can degrade when ASR transcripts are unreliable. HF Daily Papers' note
The paper replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation. It adds training-free acoustic length estimation and uses ReFlow distillation to speed the flow-matching renderer. The authors report competitive zero-shot quality across two English and two Mandarin datasets. They also test references in eight other languages, infant babble, and synthetic gibberish, arguing transcript-dependent systems can degrade when ASR transcripts are unreliable. HF Daily Papers' note
score 5