Open AI’s Astra model is on the way—and very good at breaking into computer systems
OpenAI says Astra can find and exploit unknown software flaws without human guidance.
The company calls Astra the first model to hit its “critical cybersecurity threshold” and says access to its strongest cyber features will be limited. OpenAI says the model scored perfectly on ExploitBench and found two zero-day vulnerabilities in an internal variant of the test. The company has not named its preview testers, explained how high-risk accounts are being restricted, or said whether the U.S. government is involved in evaluation. OpenAI says Astra did not try to escape its test environment in a Hugging Face-inspired experiment, but the piece notes there is no outside confirmation of the company’s safety claims. TechCrunch AI's note
The company calls Astra the first model to hit its “critical cybersecurity threshold” and says access to its strongest cyber features will be limited. OpenAI says the model scored perfectly on ExploitBench and found two zero-day vulnerabilities in an internal variant of the test. The company has not named its preview testers, explained how high-risk accounts are being restricted, or said whether the U.S. government is involved in evaluation. OpenAI says Astra did not try to escape its test environment in a Hugging Face-inspired experiment, but the piece notes there is no outside confirmation of the company’s safety claims. TechCrunch AI's note
score 8