A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem
Attacker-controlled MCP tool metadata and outputs can push agents into choosing malicious tools and following hostile returns.
The paper presents A2M, a black-box attack framework with an “Attraction” phase for improving malicious tool invocation and a “Manipulation” phase that uses execution traces to steer outcomes. On LiveMCPBench, attacks optimized on GLM-4.6 reached a 93.6% macro-average malicious tool invocation rate across four scenarios. The authors also report 32.4x weighted token costs under Cognitive Denial of Service and a 74.4% mean success rate across exfiltration, integrity compromise, and reasoning derailment scenarios. Transfer to four other models was weaker but still present, with 63.6% malicious invocation and 24.5% mean attack success. ArXiv · AI/CL/LG's note
The paper presents A2M, a black-box attack framework with an “Attraction” phase for improving malicious tool invocation and a “Manipulation” phase that uses execution traces to steer outcomes. On LiveMCPBench, attacks optimized on GLM-4.6 reached a 93.6% macro-average malicious tool invocation rate across four scenarios. The authors also report 32.4x weighted token costs under Cognitive Denial of Service and a 74.4% mean success rate across exfiltration, integrity compromise, and reasoning derailment scenarios. Transfer to four other models was weaker but still present, with 63.6% malicious invocation and 24.5% mean attack success. ArXiv · AI/CL/LG's note
score 6