Open AI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities | WIRED
Overview
Open AI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities
Open AI announced Tuesday that its forthcoming AI model, Astra, is its first to reach the company’s threshold for what it calls “critical” cyber capabilities. Open AI says it plans to publicly release a version of Astra “soon,” but will make the model’s advanced cyber capabilities available only to select partners in its Daybreak Blue early-access program at launch.
Details
In a briefing with reporters, Open AI safety and security leaders said the company has concluded that Astra reaches the critical cybersecurity capabilities outlined in its preparedness framework, which sets thresholds and protocols for when its AI models pose new levels of risk. The company says an AI model has reached its critical cyber threshold when it can independently find and exploit previously unknown vulnerabilities in real-world software. Open AI leaders said the company has followed its procedure for this situation, which is to halt further development until appropriate safeguards and security measures can be implemented.
Open AI previously said that it paused some training workloads related to the development of Astra and a future AI model for several weeks. Executives say the company has now resumed said work on Astra, and the future AI model, after putting additional safety and security controls in place. Open AI says the multi-week pause was productive, and it is now confident that it can release Astra broadly in a safe way.
The announcement comes as Silicon Valley grapples with the advanced cybersecurity capabilities of cutting-edge AI models, and tries to assure users, lawmakers, and other companies that it can keep them under control. In July, Open AI disclosed an incident in which agents running two of its models exploited vulnerabilities in what was supposed to be a siloed testing environment, gaining access to the internet and hacking the open source AI platform Hugging Face. (Open AI notes that Astra was not one of the models involved in this case.)
grapples with the advanced cybersecurity capabilities
Other AI companies, such as Anthropic and Meta, have disclosed similar incidents in recent weeks. On Monday, Anthropic also said it has paused some AI training workloads while it hardens its safety and security practices.
Open AI says it’s implementing a multi-step approach to limit everyday users from accessing Astra’s advanced cyber capabilities, including a new “misalignment monitor.” If someone asks Astra to help them find an exploit in a real-world software system, for example, the model is supposed to refuse to answer. Open AI says it has also made Astra more robust to jailbreaking attempts, and in tests it successfully refused unsafe queries at a significantly higher rate than previous models.
However, Open AI notes in a blog post that its misalignment monitor may “occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped.” Open AI says the guardrail can be triggered in some cases even when a user is engaging in activities that don’t appear related to cybersecurity. When this happens, Chat GPT and Codex users may be asked to review the model’s action before proceeding, Open AI said.
Partners in Open AI’s Daybreak program—which includes digital infrastructure providers like Cisco, Cloudflare, and Palo Alto Networks—will get early access to a less restricted version of Astra with more robust cyber capabilities. The goal of the program is to ensure these companies can use advanced AI models like Astra to harden their defenses before similarly capable models are made broadly available. Open AI leaders also said the company has been working closely with government partners to ensure they’re aware of Astra’s cyber skills and can get access to them.
Astra is not only capable of finding novel software vulnerabilities and developing ways to exploit them for hacking, but is also able to “chain” multiple exploits together, a technique used to bore deeper and deeper into a target system and gain access that wouldn’t be attainable using just one vulnerability.
According to figures from Open AI, Astra outperforms industry leading AI models such as GPT-5.6 Sol and Anthropic’s Mythos on cybersecurity benchmarks such as Exploit Bench, which Astra scored 100 percent on. However, these capabilities are broadly in line with the rising hacking abilities of AI models that Open AI and Anthropic have been forecasting for months. In April, for example, Anthropic emphasized that Mythos Preview was able to autonomously develop exploit chains.
As the AI and cybersecurity industries have scrambled to adapt, though, many cybersecurity experts have emphasized that key digital security defenses and longstanding best practices are still durable. However, AI puts organizations and systems that haven’t fully implemented these protections at even more urgent risk.
In your inbox: Maxwell Zeff's dispatch on the business of AI
In your inbox: Maxwell Zeff's dispatch on the business of AI
Big Story: How data centers broke American politics
Big Story: How data centers broke American politics
Hackers stalked me by hijacking a cheap pink plastic watch
Hackers stalked me by hijacking a cheap pink plastic watch
Special edition: The WIRED guide to expanding your mind
Special edition: The WIRED guide to expanding your mind
Key Takeaways
-
Open AI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities
-
Open AI announced Tuesday that its forthcoming AI model, Astra, is its first to reach the company’s threshold for what it calls “critical” cyber capabilities
-
In a briefing with reporters, Open AI safety and security leaders said the company has concluded that Astra reaches the critical cybersecurity capabilities outlined in its preparedness framework, which sets thresholds and protocols for when its AI models pose new levels of risk
-
Open AI previously said that it paused some training workloads related to the development of Astra and a future AI model for several weeks
-
The announcement comes as Silicon Valley grapples with the advanced cybersecurity capabilities of cutting-edge AI models, and tries to assure users, lawmakers, and other companies that it can keep them under control



