Megadose AI progress, ranked and analyzed.

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

· ArXiv · AI/CL/LG ·
The paper argues multi-hop retrieval failures can be predicted from retrieval-score structure before another LLM call is made.

Andre Bacellar defines a confidence method, RegimeAbstain, using up to nine query and ANN features to decide when to abstain. The paper says these features predict different failure regimes better than any single score signal. Across MuSiQue, 2WikiMultiHopQA, and HoVer, its Retrieval Confidence Score is reported as best or co-best on AUC-AC against eight baselines. On MuSiQue with an LLM-judge retriever, it cuts confident-wrong answers from 39.5% to 20.6% at 50% coverage. ArXiv · AI/CL/LG's note

score 5

Categories: Research