OpenAI has released GPT-6 Astra, which it calls its most intelligent and aligned model to date, and the first it has broadly deployed to reach the "Critical" cybersecurity capability threshold under its own Preparedness Framework. OpenAI reports the model scored a perfect 100% on its ExploitBench benchmark, against 78.5% for predecessor GPT-5.6 Sol.

The company is leaning heavily on alignment to justify the release. On a test built around its Hugging Face incident, which checks whether a model given an impossible task oversteps its authorised scope, OpenAI reports Astra did so in 0% of cases, against 48% for Sol. A related "honeypot" exploit test shows a similar gap: Astra made no attempts to bypass restrictions, where Sol did in nearly half.

That sits awkwardly beside OpenAI's own admission that Astra's reasoning has become harder to monitor. The model produces shorter, less informative reasoning traces, and can sometimes underperform deliberately on evaluations or slip past internal monitors under adversarial prompting. OpenAI says it found no evidence of the model hiding reasoning inside unrelated text, but calls the monitorability decline a trend it takes seriously.

At launch, Astra will refuse advanced cyber tasks such as producing proof-of-concept exploits; OpenAI plans to loosen that through its Daybreak programme for vetted users in the coming weeks. The model is rolling out now to a limited set of organisations, reaching ChatGPT Plus, Pro, Business and Enterprise users and the API within days, priced at $10 per million input tokens and $50 per million output tokens.


Execution Level Governance- What audit-ready agent governance actually looks like
David Girvin, founder and CEO of Assury argues that model-in-the-loop review, AI governing AI, is fundamentally unreliable for regulated environments: even the best-performing models miss a meaningful share of violations, the reviewing model is typically provided by the same vendor being reviewed, and prompt injection or context poisoning can compromise both the acting agent and its supposed overseer simultaneously. He makes the case for deterministic, architecturally enforced controls instead, walking through Assury’s approach of autonomy zones, session risk accumulation, and credential starvation, which lets a compromised agent be cut off from its tools instantly rather than relying on time-boxed access. The conversation touches on why David is sceptical of just-in-time credentialing as a solution for agent security more broadly, since agent sessions don’t run on predictable human timescales, along with the current gap between how identity and security vendors are pitching agent protection and what he sees happening at the execution layer in practice. He also discusses the compliance and audit implications of probabilistic decision-making, arguing that regulated industries will increasingly need tamper-evident, hash-chained audit trails that can withstand scrutiny from auditors and regulators who are only beginning to understand agentic risk, and reflects on a named frontier lab’s own published framework as an example of the gap between research and practitioner reality. Elsewhere, David reflects candidly on building a bootstrapped security company in an increasingly crowded market, why he turned down aggressive VC funding to stay in control of the product, and what a credible third-party assessment of his own gateway would need to look like given that Assury sits directly in the execution path for every customer’s agents.
Share this post
The link has been copied!