Accent Analogy Guidance: More Speaker Similarity at Equal Accent in Cross-Lingual Voice Cloning
AAG improves speaker match without paying the usual accent penalty in several cross-lingual voice-cloning models.
The paper proposes a training-free sampler term that estimates an “accent direction” from a synthetic voice rendered in two languages, then subtracts that direction during generation. In tests across four open TTS systems, AAG sits above the usual identity-accent trade-off curve for OmniVoice, MaskGCT, and CosyVoice 2; F5-TTS becomes more native-sounding than any reweighting baseline. The authors say listener checks and an LLM-free language-ID measure support the same result, while a premise test predicts the one model where AAG does not help, X-Voice. HF Daily Papers' note
The paper proposes a training-free sampler term that estimates an “accent direction” from a synthetic voice rendered in two languages, then subtracts that direction during generation. In tests across four open TTS systems, AAG sits above the usual identity-accent trade-off curve for OmniVoice, MaskGCT, and CosyVoice 2; F5-TTS becomes more native-sounding than any reweighting baseline. The authors say listener checks and an LLM-free language-ID measure support the same result, while a premise test predicts the one model where AAG does not help, X-Voice. HF Daily Papers' note
score 4