Megadose AI progress, ranked and analyzed.

To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech

· ArXiv · AI/CL/LG ·
VeriSpeak tests whether audio-language models can fact-check spoken claims with retrieved text evidence.

The benchmark contains 3,879 spoken claims across temporal, geographical, and relational facts, balanced between true and false labels. The authors report a persistent gap between written and spoken verification: models that handle text claims well often miss the same claims in speech. Retrieval by itself helps only modestly, because models can blur the evidence with the spoken claim. Adding explicit reasoning improves comparison, with a thinking-tuned audio model reaching 86.1% accuracy. ArXiv · AI/CL/LG's note

score 4

Categories: Research