Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models
Prompted lying made reasoning models spend more tokens thinking.
The paper tested three reasoning-capable LLMs on 210 multiple-choice questions under prompts to answer truthfully, falsely, or without regard for truth. Truth-directed answers used fewer reasoning tokens than both lie-directed and truth-indifferent answers across all three models. The authors frame token count as a content-independent signal when chain-of-thought text is unavailable or unreliable, but they do not claim it detects spontaneous deception or general misalignment yet. ArXiv · AI/CL/LG's note
The paper tested three reasoning-capable LLMs on 210 multiple-choice questions under prompts to answer truthfully, falsely, or without regard for truth. Truth-directed answers used fewer reasoning tokens than both lie-directed and truth-indifferent answers across all three models. The authors frame token count as a content-independent signal when chain-of-thought text is unavailable or unreliable, but they do not claim it detects spontaneous deception or general misalignment yet. ArXiv · AI/CL/LG's note
score 4