Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors
Richer multi-hop RAG kept higher scores, but made ASR mistakes hurt more.
The paper tests entity-graph linking and iterative reformulation on spoken-query pipelines using four synthesized English accents and three multi-hop QA benchmarks. Under ASR input, the combined richer setup showed a clean-to-highest-WER F1 gap 36-67% larger than naive dense retrieval across all three benchmarks. Corrupted query entities were the main failure mode, making up 87-96% of degradation cases on 2WikiMultiHopQA. Lightweight surface-form fixes left most of the gap unresolved. HF Daily Papers' note
The paper tests entity-graph linking and iterative reformulation on spoken-query pipelines using four synthesized English accents and three multi-hop QA benchmarks. Under ASR input, the combined richer setup showed a clean-to-highest-WER F1 gap 36-67% larger than naive dense retrieval across all three benchmarks. Corrupted query entities were the main failure mode, making up 87-96% of degradation cases on 2WikiMultiHopQA. Lightweight surface-form fixes left most of the gap unresolved. HF Daily Papers' note
score 4