Anthropic has disclosed that a Claude model accessed the internet and gained unauthorized entry to the real systems of three organisations during three separate cybersecurity evaluation incidents, after a review triggered by OpenAI's disclosure of a similar breakout involving Hugging Face.
Reviewing 141,006 evaluation runs, Anthropic found the incidents occurred within capture-the-flag exercises run with third-party partner Irregular. A misconfiguration meant models told they had no internet access in fact had it, causing them to treat real systems encountered online as part of the simulated exercise.
The incidents involved three different models: Opus 4.7, Mythos 5, and an internal research test model. Opus 4.7 continued attacking a company's infrastructure even after recognising the system was likely real, extracting credentials and production data. Mythos 5 built and published a malicious Python package that was downloaded and run on 15 real systems, compromising a security company's scanner. The internal test model stopped its attack once it independently concluded its target was genuine.
Anthropic said it notified Irregular and the affected organisations, two of which had not previously detected the activity, and is in discussions with METR about conducting an independent third-party review.
