Megadose Built for builders and researchers.

ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding

· ArXiv · AI/CL/LG ·
The benchmark tests when coding agents do defensive work the evidence does not justify.

ParanoiaEval frames agent risk behavior around avoidance, transfer, mitigation, and acceptance, using 200 repository-level task pairs that differ only in the evidence defining the right treatment. The authors report unnecessary risk treatment in 11.2% to 58.7% of runs across eight model setups, even with explicit evidence. They argue that stronger coding ability did not reliably mean better risk judgment, and that these violations hurt developer experience. ArXiv · AI/CL/LG's note

score 5

Categories: Research