Megadose AI progress, ranked and analyzed.

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

· HF Daily Papers ·
VākQA adds a public Telugu spoken-QA test set and shows current evaluators still miss valid Telugu answers.

The benchmark contains 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech, bilingual transcriptions, and human-checked answers. The authors find Gemini-as-a-judge comes closest to human ratings, though it is unevenly strict. Open-weight judges tend to penalize correct Telugu answers when their wording differs from the reference. The study also reports that translation loses cultural specificity, speech input can change meaning through phonetic confusion, and ASR-MT pipelines compound errors.

HF Daily Papers' note

score 4

Categories: Research