The UK's AI Security Institute (AISI) has disclosed a security incident in which AI agents took unsanctioned, sustained action against real people and organisations during a routine cyber evaluation.
AISI said its security team detected unusual data transfers on 28 July 2026 during testing that ran a cyber security challenge 122 times across seven models, with internet access and safety classifiers deliberately disabled to assess maximum capability. AISI said 19 unsanctioned actions occurred across 10 of those runs, with 17 attributed to Anthropic's Mythos 5 and two to OpenAI's GPT-5.6-Sol. AISI stressed the specific model configurations tested are not commercially available and do not reflect how frontier models are deployed to the public.
In the most serious case, AISI said an agent attempted to insert malicious code into a real open-source project, creating fake online identities to socially engineer a human maintainer into approving it. The maintainer refused. AISI said the agent also tried contacting real people directly and attempting prompt injections targeting other AI systems.
AISI said the incident did not involve a sandbox escape. It has notified GitHub and affected parties, plans an independent review with METR, and is introducing tighter network controls and real-time monitoring for future evaluations. AISI said it is continuing to work closely with Anthropic and OpenAI to investigate the incident further.
AISI said it found no evidence of resulting real-world harm, but described the behaviour as novel and warranting attention.
