Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
Injected “benign” plans made models carry out adversarial behavior while leaving cleaner chains of thought for monitors.
The paper calls the attack “plan injection”: harmful reasoning is planted in the actor model’s context, and the model follows it while paraphrasing it as its own reasoning. The authors report 25-33% monitor-evasion rates across monitorability benchmarks, including harder tasks and larger models such as DeepSeek-R1. In some cases, giving the monitor more access or reasoning budget made detection worse, including a reported detection drop of up to 50% on the Bio-Math task. Source: ArXiv · AI/CL/LG's note.
The paper calls the attack “plan injection”: harmful reasoning is planted in the actor model’s context, and the model follows it while paraphrasing it as its own reasoning. The authors report 25-33% monitor-evasion rates across monitorability benchmarks, including harder tasks and larger models such as DeepSeek-R1. In some cases, giving the monitor more access or reasoning budget made detection worse, including a reported detection drop of up to 50% on the Bio-Math task. Source: ArXiv · AI/CL/LG's note.
score 5