Humans missed 1 in 3 threats approving AI agent commands across 40k game runs

· HN · Agents ·

A large-scale study found humans often failed to catch dangerous AI agent commands, highlighting limits of manual oversight.

Categories: Research

Excerpt

HN · 116 points · 85 comments

Discussions