ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
The benchmark tests when coding agents do defensive work the evidence does not justify.
ParanoiaEval frames agent risk behavior around avoidance, transfer, mitigation, and acceptance, using 200 repository-level task pairs that differ only in the evidence defining the right treatment. The authors report unnecessary risk treatment in 11.2% to 58.7% of runs across eight model setups, even with explicit evidence. They argue that stronger coding ability did not reliably mean better risk judgment, and that these violations hurt developer experience. ArXiv · AI/CL/LG's note
ParanoiaEval frames agent risk behavior around avoidance, transfer, mitigation, and acceptance, using 200 repository-level task pairs that differ only in the evidence defining the right treatment. The authors report unnecessary risk treatment in 11.2% to 58.7% of runs across eight model setups, even with explicit evidence. They argue that stronger coding ability did not reliably mean better risk judgment, and that these violations hurt developer experience. ArXiv · AI/CL/LG's note
score 5