OpenAI has stated that it cannot rule out that an upcoming model, codenamed Astra, has reached "critical" cybersecurity capability under its Preparedness Framework, following internal evaluations conducted over the preceding days.
The company said expert assessments and recent test results led it to this conclusion on the night before publication. Under the framework, a model crosses the critical threshold if it can identify and develop functional zero-day exploits across all severity levels in hardened real-world systems without human help, or independently devise and execute end-to-end cyberattack strategies against hardened targets from only a high-level goal.
Previous models, including GPT-5.6 Sol, had been assessed at the lower "High" threshold. OpenAI said Astra was not involved in the earlier Hugging Face security incident.
In response, OpenAI said it has scaled up robustness testing of its safeguards, imposed stricter security controls on higher-capability models including isolated testing environments and enhanced weight protections, paused internal work on Astra that does not yet meet those controls, and introduced monitoring of the model's chain of thought to interrupt high-risk activity. It also plans to work with government agencies and AI safety organisations to test the model's capabilities.
OpenAI said it was applying the same precautionary approach it used in June 2025 when its models approached the high biological capability threshold.
