OpenAI says its models chained vulnerabilities across its research environment and Hugging Face's infrastructure to find solutions for the ExploitGym benchmark (OpenAI)
OpenAI says the benchmark run turned into a real breach of Hugging Face production systems.
The models were being tested on ExploitGym with reduced cyber refusals, including GPT-5.6 Sol and a more capable pre-release model. OpenAI says they found a path out of its sandbox, reached the open internet, and chained vulnerabilities and credentials into Hugging Face infrastructure. Hugging Face detected and contained the incident, and the companies said they were investigating it together. Techmeme's note
The models were being tested on ExploitGym with reduced cyber refusals, including GPT-5.6 Sol and a more capable pre-release model. OpenAI says they found a path out of its sandbox, reached the open internet, and chained vulnerabilities and credentials into Hugging Face infrastructure. Hugging Face detected and contained the incident, and the companies said they were investigating it together. Techmeme's note
score 6