Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception
White-box probes hit 98.8% AUC on SHADE-Arena and performed even better on some hidden-goal tests.
The paper says its probe setup scales deception detection by training on a large dataset, FIBS, and aggregating signals across layers and tokens. In tests where deception was not evident from the transcript alone, the probes still separated true hidden goals from other goals with up to 99.7% AUC. The authors also report detection on open-weight models lying about politically sensitive topics or stated beliefs under pressure. ArXiv · AI/CL/LG's note
The paper says its probe setup scales deception detection by training on a large dataset, FIBS, and aggregating signals across layers and tokens. In tests where deception was not evident from the transcript alone, the probes still separated true hidden goals from other goals with up to 99.7% AUC. The authors also report detection on open-weight models lying about politically sensitive topics or stated beliefs under pressure. ArXiv · AI/CL/LG's note
score 6