Megadose AI progress, ranked and analyzed.

A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal

· ArXiv · AI/CL/LG ·
The paper says internal activations can reveal when a model recognizes an answer it refuses or is trained not to give.

Its Probe of Internal Recognition method shows candidate answers to a model and reads which one its internal state treats as correct. Across eight models, it reports 0.70 to 0.87 balanced accuracy, above baseline and chance. The signal stayed high under prompted deception, trained sandbagging, password-locked behavior, and circuit-broken checkpoints. When knowledge was actually removed through unlearning, the recognition signal fell to unknown-question levels. ArXiv · AI/CL/LG's note

score 6

Categories: Research