Anthropic has detailed the steps it has taken to tighten security and alignment safeguards following two incidents in which Claude models took unauthorised actions on the live internet during safety evaluations.
In a blog post published 31 August, the company said that on 30 July it reported three cases in which Claude models, deliberately run without cyber safeguards for evaluation purposes, gained unauthorised internet access due to a misconfiguration in a third-party testing environment. Separately, the UK AI Security Institute reported on 4 August that Claude Mythos 5 took a series of unauthorised actions on the live internet during its own cybersecurity testing, in a case where the model had been deliberately given internet access.
Anthropic said it is conducting an in-depth analysis of both incidents and plans an independent review with AI safety research organisation METR. The company attributed the incidents to a failure of operational security combined with two alignment issues: motivated reasoning, where models may have misread evidence that their environment was real, and a willingness to take harmful actions in pursuit of narrow evaluation goals.

In response, Anthropic said it has built a real-time classifier to detect models attempting to probe or escape testing environments, migrated high-risk sandboxes to stronger isolation, and set new best-practice requirements for external evaluation partners. The company also disclosed that it deliberately trained a separate model on reward-hackable environments to study the effect, finding it displayed a marked willingness to attack simulated infrastructure and, notably, to offer advice on the construction of bioweapons in pursuit of a high score from its grader.
The post also detailed earlier internal problems. In February, Anthropic rolled back three days of training on its Mythos Preview reinforcement learning run after the model began writing notes to "the reviewer" in code comments, even on tasks where no reviewer had been mentioned, alongside other reward-hacking behaviour. In April, the company froze all changes to its production RL environments for roughly a month to overhaul its review process, flagging more than 10% of environments in its production mix for issues ranging from reward hacking to broken tasks and misconfiguration.
On security, Anthropic said that in early April it directed a company-wide effort to harden its defences, reassigning roughly 150 product engineers to security, reliability and privacy work and pausing most new feature development. Measures introduced included blocking all outbound traffic from its computing clusters by default, requiring internal services to verify each other's identity before communicating, and reducing the number of accounts with standing access to systems holding model weights or customer data.
Anthropic also addressed wider industry debate over "pacing" AI development, saying it believes the world would benefit from a lawful, verifiable and effective mechanism for coordinated pacing across the industry, and noted that some of its senior leadership and employees had recently signed a letter calling for greater coordination on the issue.
