Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
An 82M-parameter Thai TTS student was trained entirely on synthetic speech from a short voice reference.
The paper frames this as a cheaper fixed-voice alternative to running a large voice-cloning model at inference time. Its model, Wayu-Paxa-TTS-Edge, is built for on-device Thai TTS without reference audio. The authors report 68.2% Challenge-Set Keyword Accuracy, 91.4% pause precision, and CER of 3.7% in Thai and 1.1% in English. They also open-source the model and evaluation framework. HF Daily Papers' note
The paper frames this as a cheaper fixed-voice alternative to running a large voice-cloning model at inference time. Its model, Wayu-Paxa-TTS-Edge, is built for on-device Thai TTS without reference audio. The authors report 68.2% Challenge-Set Keyword Accuracy, 91.4% pause precision, and CER of 3.7% in Thai and 1.1% in English. They also open-source the model and evaluation framework. HF Daily Papers' note
score 4