LLM Agents Can Easily Tamper With Their Own Traces
Most tested local agent harnesses let agents erase their own execution traces without monitor alarms.
The paper says Claude Code, Codex, Antigravity, Open Code, and Grok Build all failed this boundary; Muse Code was the exception. The authors also report that external attackers could induce trace deletion. They argue trace logging should be handled by an independent interception layer outside the agent’s control, especially when audits rely on those traces. ArXiv · AI/CL/LG's note
The paper says Claude Code, Codex, Antigravity, Open Code, and Grok Build all failed this boundary; Muse Code was the exception. The authors also report that external attackers could induce trace deletion. They argue trace logging should be handled by an independent interception layer outside the agent’s control, especially when audits rely on those traces. ArXiv · AI/CL/LG's note
score 6