Megadose AI progress, ranked and analyzed.

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal

· HF Daily Papers ·
PIR reads internal recognition to tell hidden knowledge from missing knowledge.

The paper adapts the Concealed Information Test to language models by showing candidate answers and measuring which one the model’s internal states recognize as correct. Across eight models, it reports 0.70 to 0.87 balanced accuracy, above unknown-item baselines and chance. The method still found recognition under prompted deception, trained sandbagging, password-locked behavior, and circuit-broken checkpoints. When unlearning actually removed the knowledge, the recognition signal fell to the level of never-known questions. HF Daily Papers' note

score 5

Categories: Research