Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
Google says frontier models usually store the fact; the failure is getting it back out.
The post introduces “knowledge profiling,” tested through WikiProfile, a benchmark of 2,150 Wikipedia-derived facts. In Gemini-3-Pro and GPT-5, Google reports 95-98% of facts were encoded, but 26-34% still could not be directly recalled. Thinking reduced the miss rate, but 11-12% of encoded facts still failed. The authors frame rare facts and reverse questions as recall problems, not simply missing knowledge. Google Research's note
The post introduces “knowledge profiling,” tested through WikiProfile, a benchmark of 2,150 Wikipedia-derived facts. In Gemini-3-Pro and GPT-5, Google reports 95-98% of facts were encoded, but 26-34% still could not be directly recalled. Thinking reduced the miss rate, but 11-12% of encoded facts still failed. The authors frame rare facts and reverse questions as recall problems, not simply missing knowledge. Google Research's note
score 5