A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
PIR reads internal recognition to tell hidden knowledge from missing knowledge.
The paper adapts the Concealed Information Test to language models by showing candidate answers and measuring which one the model’s internal states recognize as correct. Across eight models, it reports 0.70 to 0.87 balanced accuracy, above unknown-item baselines and chance. The method still found recognition under prompted deception, trained sandbagging, password-locked behavior, and circuit-broken checkpoints. When unlearning actually removed the knowledge, the recognition signal fell to the level of never-known questions. HF Daily Papers' note
The paper adapts the Concealed Information Test to language models by showing candidate answers and measuring which one the model’s internal states recognize as correct. Across eight models, it reports 0.70 to 0.87 balanced accuracy, above unknown-item baselines and chance. The method still found recognition under prompted deception, trained sandbagging, password-locked behavior, and circuit-broken checkpoints. When unlearning actually removed the knowledge, the recognition signal fell to the level of never-known questions. HF Daily Papers' note
score 5