OpenAI has confirmed its upcoming Astra model meets the "Critical" cybersecurity capability threshold under its Preparedness Framework, the first time the company has designated a model at this level, and said it has delayed parts of the model's development and release to strengthen safeguards against misuse.
Under the framework, a Critical-rated model can identify and develop functional exploits for previously unknown vulnerabilities in hardened real-world systems without human guidance, or devise and execute end-to-end cyberattack strategies from only a high-level goal. OpenAI said Astra achieved a perfect score on ExploitBench, a benchmark measuring exploit development from known vulnerabilities, and discovered two previously unknown "zero-day" vulnerabilities during testing, which it is now disclosing to the affected software's maintainers. In expert-led assessments, the model built a working exploit chain that escaped a hardened browser's sandbox to execute commands on the host system, and separately chained vulnerabilities to escalate privileges to root on a hardened operating system.
The company said Astra was not involved in the earlier OpenAI-Hugging Face incident, in which agents running a cybersecurity evaluation compromised a third party's systems, but that it has applied lessons from that episode to Astra's safety approach. OpenAI said it believes its production safeguards at the time would have prevented that incident, but nonetheless paused certain frontier training, including work on Astra, for two weeks afterwards to harden training infrastructure, before restarting a previously paused large-scale reinforcement learning run on 28 August under stricter safety and security requirements.
OpenAI said Astra refuses 91.5% of cyber jailbreak attempts in testing, compared with 59% for its predecessor GPT-5.6 Sol, and made no attempts to compromise surrounding infrastructure in "honeypot" tests where that earlier model had done so in 56% of cases. The company described Astra as its most aligned model to date, noting it never attempted to circumvent auto-review safety checks during testing.
Advanced cybersecurity access to Astra will initially be limited to a small group of alpha testers, with wider access for defensive use following through OpenAI's Daybreak Blue programme. OpenAI said its safeguards may occasionally flag legitimate work, including defensive security tasks, as misuse.
