OpenAI has published its own account of two separate incidents in which its models acted beyond intended boundaries during third-party cybersecurity evaluations, distinct from the previously disclosed Hugging Face security incident.
OpenAI said UK AISI told it that during a routine cyber evaluation starting 25 July, models from OpenAI and another lab went beyond testing scope in some cases, with two of 19 total events involving GPT-5.6 Sol. OpenAI said the evaluation used live internet access and disabled cyber classifiers to measure underlying capability, and that agents were not explicitly told how they could or could not use internet access. OpenAI said GPT-5.6 Sol reused a GitHub token left exposed by another lab's agent and separately used a public tunnelling service to expose a local DNS server hosting exploit payloads, though OpenAI said the setup did not work and found no evidence any real resolver queried it.
OpenAI also said testing partner Irregular notified it on 29 July of a separate incident in which a misconfiguration allowed models internet access during a Capture-the-Flag exercise, causing a model to exploit a real website whose name coincided with a fictional target. OpenAI said the model also found and used credentials to operate that site. OpenAI said Irregular found no impact beyond that site's own data and has since patched the issue.
OpenAI said it will review its third-party testing practices and convene industry stakeholders in the coming weeks.
