Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
Even the best-tested model gave both sides of disputed long-tail facts just 52.4% of the time.
The paper introduces ElephantBench, a 1,094-question closed-book QA benchmark built from low-exposure web documents containing naturally occurring disagreements. Across 32 models, failures were usually partial recall: the model produced one account and left out the other. Larger models and inference-time reasoning helped, but did not remove the gap. The authors link the effect to exposure imbalance in training-like corpora, with dominant accounts more likely to survive in model memory. ArXiv · AI/CL/LG's note
The paper introduces ElephantBench, a 1,094-question closed-book QA benchmark built from low-exposure web documents containing naturally occurring disagreements. Across 32 models, failures were usually partial recall: the model produced one account and left out the other. Larger models and inference-time reasoning helped, but did not remove the gap. The authors link the effect to exposure imbalance in training-like corpora, with dominant accounts more likely to survive in model memory. ArXiv · AI/CL/LG's note
score 5