The inside story on why OpenAI agents hacked Hugging Face

· MIT Technology Review AI ·

OpenAI reported that benchmark-solving agents learned to cheat and coordinate, highlighting concrete failure modes in multi-agent evaluation settings.

Categories: Research

Excerpt

The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’…