OpenAI has confirmed its upcoming Astra model meets the "Critical" cybersecurity capability threshold under its Preparedness Framework, the first time the company has designated a model at this level, and said it has delayed parts of the model's development and release to strengthen safeguards against misuse.

Under the framework, a Critical-rated model can identify and develop functional exploits for previously unknown vulnerabilities in hardened real-world systems without human guidance, or devise and execute end-to-end cyberattack strategies from only a high-level goal. OpenAI said Astra achieved a perfect score on ExploitBench, a benchmark measuring exploit development from known vulnerabilities, and discovered two previously unknown "zero-day" vulnerabilities during testing, which it is now disclosing to the affected software's maintainers. In expert-led assessments, the model built a working exploit chain that escaped a hardened browser's sandbox to execute commands on the host system, and separately chained vulnerabilities to escalate privileges to root on a hardened operating system.

The company said Astra was not involved in the earlier OpenAI-Hugging Face incident, in which agents running a cybersecurity evaluation compromised a third party's systems, but that it has applied lessons from that episode to Astra's safety approach. OpenAI said it believes its production safeguards at the time would have prevented that incident, but nonetheless paused certain frontier training, including work on Astra, for two weeks afterwards to harden training infrastructure, before restarting a previously paused large-scale reinforcement learning run on 28 August under stricter safety and security requirements.

OpenAI said Astra refuses 91.5% of cyber jailbreak attempts in testing, compared with 59% for its predecessor GPT-5.6 Sol, and made no attempts to compromise surrounding infrastructure in "honeypot" tests where that earlier model had done so in 56% of cases. The company described Astra as its most aligned model to date, noting it never attempted to circumvent auto-review safety checks during testing.

Advanced cybersecurity access to Astra will initially be limited to a small group of alpha testers, with wider access for defensive use following through OpenAI's Daybreak Blue programme. OpenAI said its safeguards may occasionally flag legitimate work, including defensive security tasks, as misuse.


Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.
Share this post
The link has been copied!