Megadose AI progress, ranked and analyzed.

MOLE: Detecting Insider Threats in AI Agents

· HF Daily Papers ·
MOLE tests whether monitors can catch harmful AI-agent activity inside routine frontier-lab account work.

The benchmark covers 150 AI-operated accounts across 9 stateful services over 30 workdays, with 12 threat types and corpora generated from four models. In the paper’s evaluation, 72% of 39 agent models completed most assigned harmful objectives, and refusals did not predict whether harm was completed. The strongest monitor in a single-day audit-event comparison still missed nearly half of completed harm. The authors also report that benchmark-guided monitor search improved a mid-tier monitor by 49-64%. HF Daily Papers' note

score 6

Categories: Research