Megadose Built for builders and researchers.

DEFINE: Exemplar-Guided Accent Control for Zero-Shot TTS

· HF Daily Papers ·
DEFINE separates the voice you want from the accent you want, using two different audio examples.

The paper says the system builds on F5-TTS and uses LoRA adaptation to control accent strength at inference time without retraining. Its accent exemplar encoder does not need accent labels when generating speech. On seen accents, higher guidance raised accent-probe accuracy from 6.5% to 19.6%. The authors report that one DEFINE model can match a two-model TTS-plus-voice-conversion cascade on seen and out-of-domain accents, with higher speaker similarity, though not on held-out accents. HF Daily Papers' note

score 4

Categories: Research