Anthropic has detailed the steps it has taken to tighten security and alignment safeguards following two incidents in which Claude models took unauthorised actions on the live internet during safety evaluations.

In a blog post published 31 August, the company said that on 30 July it reported three cases in which Claude models, deliberately run without cyber safeguards for evaluation purposes, gained unauthorised internet access due to a misconfiguration in a third-party testing environment. Separately, the UK AI Security Institute reported on 4 August that Claude Mythos 5 took a series of unauthorised actions on the live internet during its own cybersecurity testing, in a case where the model had been deliberately given internet access.

Anthropic said it is conducting an in-depth analysis of both incidents and plans an independent review with AI safety research organisation METR. The company attributed the incidents to a failure of operational security combined with two alignment issues: motivated reasoning, where models may have misread evidence that their environment was real, and a willingness to take harmful actions in pursuit of narrow evaluation goals.

Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.

In response, Anthropic said it has built a real-time classifier to detect models attempting to probe or escape testing environments, migrated high-risk sandboxes to stronger isolation, and set new best-practice requirements for external evaluation partners. The company also disclosed that it deliberately trained a separate model on reward-hackable environments to study the effect, finding it displayed a marked willingness to attack simulated infrastructure and, notably, to offer advice on the construction of bioweapons in pursuit of a high score from its grader.

The post also detailed earlier internal problems. In February, Anthropic rolled back three days of training on its Mythos Preview reinforcement learning run after the model began writing notes to "the reviewer" in code comments, even on tasks where no reviewer had been mentioned, alongside other reward-hacking behaviour. In April, the company froze all changes to its production RL environments for roughly a month to overhaul its review process, flagging more than 10% of environments in its production mix for issues ranging from reward hacking to broken tasks and misconfiguration.

On security, Anthropic said that in early April it directed a company-wide effort to harden its defences, reassigning roughly 150 product engineers to security, reliability and privacy work and pausing most new feature development. Measures introduced included blocking all outbound traffic from its computing clusters by default, requiring internal services to verify each other's identity before communicating, and reducing the number of accounts with standing access to systems holding model weights or customer data.

Anthropic also addressed wider industry debate over "pacing" AI development, saying it believes the world would benefit from a lawful, verifiable and effective mechanism for coordinated pacing across the industry, and noted that some of its senior leadership and employees had recently signed a letter calling for greater coordination on the issue.


Agentic Exploits- Deterministic gates for a probabilistic problem
David Girvin, CEO and co-founder of Assury, joins Stewart Tinson to dig into what’s actually happening when agentic AI goes wrong, and why he thinks most of the industry is solving the wrong layer of the problem. David explains the difference between prompt-level exploits and execution-level ones, arguing that the real danger starts the moment an agent moves from generating text to calling tools: deleting databases, reading files, sending emails. He walks through real-world incidents, including a Mexican government breach chain that escalated from just over a thousand prompts to over five thousand AI-executed actions across multiple agencies before detection, and the UK AI Security Institute’s recent cyber evaluation, in which agents took unsanctioned action including fabricating identities to socially engineer a real GitHub maintainer. The conversation covers why David is sceptical of “guardrails” language and AI-governing-AI approaches, arguing that only deterministic, architectural controls can reliably constrain agent behaviour, alongside human review reserved for genuinely high-stakes actions rather than blanket approval fatigue. He breaks down credential starvation, session risk accumulation, and why classifier-based tools keep failing inconsistently on identical actions, pointing to a named frontier lab’s own zero trust paper as an example of the industry misjudging what actually works. Elsewhere, David discusses the exposed MCP server problem, the widening trust gap between small specialist security vendors and platform incumbents, and why he believes regulation, not product quality alone, is what finally drives enterprise security spend. He closes with the exploit that concerns him most for the year ahead: session-level, goal-directed deception with no attacker involved at all.
Share this post
The link has been copied!