Megadose AI progress, ranked and analyzed.

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

· ArXiv · AI/CL/LG ·
The paper argues that deceptive-looking model behavior is not enough to prove a deceptive mechanism, much less agency.

Yakov Pyotr Shkolnikov proposes a causal taxonomy for separating outputs that look deceptive from the internal or objective-driven processes that could produce them. The paper tests those distinctions in controlled guessing-game and stock-trading experiments on two open-weight model families. It reports cases where deceptive-looking behavior appears without the proposed mechanism, and other cases where a recipient’s information state causally affects deceptive preference. ArXiv · AI/CL/LG's note

score 4

Categories: Research