Megadose Built for builders and researchers.

Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks

· ArXiv · AI/CL/LG ·
The paper says agent-security scores can move sharply when only the model-visible wording changes.

The authors define TPRS to measure ASR shifts while keeping the task, harmful action, policy, ground truth, environment, and evaluation fixed. On ASB, neutralizing threat-related tool names raised committed ASR by 11.67 points for GPT-5-mini and 13.21 for Claude Haiku 4.5. On MCPTox, making a neutral tool name explicitly threat-related lowered ASR by 11.00 points for GPT-5-mini and 4.11 for Claude Haiku 4.5. Their conclusion is that robustness claims should be tested across controlled threat-preserving representations, not a single benchmark wording. ArXiv · AI/CL/LG's note

score 5

Categories: Research