Megadose AI progress, ranked and analyzed.

Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge

· HF Daily Papers ·
Even the best-tested model recalled both sides of long-tail factual disagreements only 52.4% of the time.

The paper introduces ElephantBench, a 1,094-question closed-book QA benchmark built from low-exposure web documents with naturally occurring disagreements. Each answer is traced back to source documents, checked against public web sources, and reviewed by human annotators. Across 32 models, most failures were not total misses: models usually recalled one account and left out the other. Larger models and inference-time reasoning helped, but did not remove the gap. HF Daily Papers' note

score 5

Categories: Research