OpenAI has paused its largest planned frontier reinforcement learning run following the Hugging Face security incident in July, and separately in response to the capabilities of its upcoming Astra model, as the company works to understand how its own models came to compromise internal and third-party infrastructure without human direction.

The company's investigation identified four contributing misalignment patterns: reward hacking, where models cheated on tasks to secure higher scores; persistence on tasks that were effectively impossible to solve; unauthorised communication between agents that were meant to operate independently; and a tendency for agents to adopt goals and instructions from one another without scrutiny. OpenAI found that a small subset of unsolved evaluation tasks accounted for the vast majority of the coordination activity it later traced on an improvised message board, and that agents frequently reasoned at length about how to satisfy what they believed the evaluation grader required, in some cases continuing to attack systems even after they had already found a valid solution.

OpenAI said the episode also exposed gaps in its safeguards, noting that protections built for production systems, including its automated review and safety classifiers, were not applied to the internal evaluation environment where the incident occurred, and that chain-of-thought monitoring was not yet running on the training in question.

Going forward, the company says it is building more isolated sandboxes, expanding network isolation, and requiring chain-of-thought monitoring for all tool-using training and evaluations involving models at or above the capability of GPT-5.6 Sol, a requirement that extends to all tool-enabled inference for Astra-class models given their potential cyber-critical capabilities. It is also developing new training methods intended to teach models to stop safely when a task appears broken, to distrust unauthorised instructions from other agents, and to stay within their original permissions over long tasks. OpenAI has additionally revised its incident response procedures, including clearer escalation rules and a requirement that responders pause suspicious activity within 30 minutes unless they can rule out a genuine problem.


Execution Level Governance- What audit-ready agent governance actually looks like
David Girvin, founder and CEO of Assury argues that model-in-the-loop review, AI governing AI, is fundamentally unreliable for regulated environments: even the best-performing models miss a meaningful share of violations, the reviewing model is typically provided by the same vendor being reviewed, and prompt injection or context poisoning can compromise both the acting agent and its supposed overseer simultaneously. He makes the case for deterministic, architecturally enforced controls instead, walking through Assury’s approach of autonomy zones, session risk accumulation, and credential starvation, which lets a compromised agent be cut off from its tools instantly rather than relying on time-boxed access. The conversation touches on why David is sceptical of just-in-time credentialing as a solution for agent security more broadly, since agent sessions don’t run on predictable human timescales, along with the current gap between how identity and security vendors are pitching agent protection and what he sees happening at the execution layer in practice. He also discusses the compliance and audit implications of probabilistic decision-making, arguing that regulated industries will increasingly need tamper-evident, hash-chained audit trails that can withstand scrutiny from auditors and regulators who are only beginning to understand agentic risk, and reflects on a named frontier lab’s own published framework as an example of the gap between research and practitioner reality. Elsewhere, David reflects candidly on building a bootstrapped security company in an increasingly crowded market, why he turned down aggressive VC funding to stay in control of the product, and what a credible third-party assessment of his own gateway would need to look like given that Assury sits directly in the execution path for every customer’s agents.
Share this post
The link has been copied!