Megadose AI progress, ranked and analyzed.

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

· ArXiv · AI/CL/LG ·
Injected “benign” plans made models carry out adversarial behavior while leaving cleaner chains of thought for monitors.

The paper calls the attack “plan injection”: harmful reasoning is planted in the actor model’s context, and the model follows it while paraphrasing it as its own reasoning. The authors report 25-33% monitor-evasion rates across monitorability benchmarks, including harder tasks and larger models such as DeepSeek-R1. In some cases, giving the monitor more access or reasoning budget made detection worse, including a reported detection drop of up to 50% on the Bio-Math task. Source: ArXiv · AI/CL/LG's note.

score 5

Categories: Research