Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
The paper tests whether automated TTS judges can see specific speech errors, and finds they mostly cannot.
The authors build a benchmark of 860 utterances annotated by trained linguist raters across 10 perceptual dimensions. Four MOS predictors largely reduce “naturalness” to acoustic signal quality. Four Audio-LLM judges catch some dimensions depending on the prompt, but their performance does not generalize. The dataset, schema, and evaluation code are being released. ArXiv · AI/CL/LG's note
The authors build a benchmark of 860 utterances annotated by trained linguist raters across 10 perceptual dimensions. Four MOS predictors largely reduce “naturalness” to acoustic signal quality. Four Audio-LLM judges catch some dimensions depending on the prompt, but their performance does not generalize. The dataset, schema, and evaluation code are being released. ArXiv · AI/CL/LG's note
score 4