Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA
Index updates changed answers even when the RAG system’s visible settings stayed fixed.
The paper names this “accuracy-blind answer churn”: answer changes that aggregate accuracy can miss when gains and losses offset each other. Its Snapshot Compatibility Audit compares repeat disagreement within the same snapshot against disagreement across expanded index snapshots. In a 400-question Natural Questions study, semantic excess churn was 10.25 percentage points while exact-match accuracy moved only -1.50 points. A smaller TriviaQA study and a second-generator subset replication showed the same direction of effect. HF Daily Papers' note
The paper names this “accuracy-blind answer churn”: answer changes that aggregate accuracy can miss when gains and losses offset each other. Its Snapshot Compatibility Audit compares repeat disagreement within the same snapshot against disagreement across expanded index snapshots. In a 400-question Natural Questions study, semantic excess churn was 10.25 percentage points while exact-match accuracy moved only -1.50 points. A smaller TriviaQA study and a second-generator subset replication showed the same direction of effect. HF Daily Papers' note
score 4