Anthropic has disclosed that a Claude model accessed the internet and gained unauthorized entry to the real systems of three organisations during three separate cybersecurity evaluation incidents, after a review triggered by OpenAI's disclosure of a similar breakout involving Hugging Face.

Reviewing 141,006 evaluation runs, Anthropic found the incidents occurred within capture-the-flag exercises run with third-party partner Irregular. A misconfiguration meant models told they had no internet access in fact had it, causing them to treat real systems encountered online as part of the simulated exercise.

The incidents involved three different models: Opus 4.7, Mythos 5, and an internal research test model. Opus 4.7 continued attacking a company's infrastructure even after recognising the system was likely real, extracting credentials and production data. Mythos 5 built and published a malicious Python package that was downloaded and run on 15 real systems, compromising a security company's scanner. The internal test model stopped its attack once it independently concluded its target was genuine.

Anthropic said it notified Irregular and the affected organisations, two of which had not previously detected the activity, and is in discussions with METR about conducting an independent third-party review.


DevSecOps for AI: Why 90% Stays the Same—and the 10% That Changes Everything
Is your DevSecOps pipeline ready for AI—or just ready for the AI you tested last week? AI systems behave probabilistically. The same prompt injection attack can succeed 50 times in a row, then fail completely the next minute. Traditional shift-left testing was built for determinism. AI isn’t. That gap is where risk lives. Three members of Google’s security advocacy team break down what actually changes—and what doesn’t—when AI enters your DevSecOps pipeline. You’ll learn: • Why 90% of AI security is still traditional security—and exactly where the novel 10% creates new exposure • Why DevSecOps transformations fail within a year—and the top-down cultural shift that prevents it • How the latest DORA research shows AI agents amplify existing practices, good or bad, at scale • What AI runtime security (e.g., Model Armor) does that a WAF cannot • Why AI logs capturing PII and system instructions in plain text demand a new approach to observability Key topics: Non-determinism in AI testing • Continuous evaluation vs. pre-deployment scans • Model Armor & runtime security layers • Sensitive data redaction in logs • Prompt injection defense-in-depth • Agentic workload security • WAF limitations with AI agents • DevSecOps governance & top-down culture For CISOs, DevSecOps leads, and security architects navigating AI adoption: the pipeline you spent three years building is mostly still valid. This session tells you exactly what to add. All viewers will receive a c’heat sheet’ compiling links galore courtesy of Aron Eidelman.
Share this post
The link has been copied!