Megadose Built for builders and researchers.

Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark

· ArXiv · AI/CL/LG ·
The paper says HarmBench does not support a single “harmful refusal” score.

The authors audit HELM Safety datasets that might measure refusal of dangerous prompts and find most are already saturated. On HarmBench, psychometric tests suggest the benchmark mixes multiple behaviors rather than isolating one attribute. They also find developer-linked item differences at the same refusal-ability level, though many fade when matching by narrower scope. The paper argues safety aggregates should prove they measure one thing before being used to compare models. ArXiv · AI/CL/LG's note

score 4

Categories: Research