Summary
- OpenAI says additional testing shows Astra meets the Critical cybersecurity threshold in its Preparedness Framework.
- The company says the model can identify previously unknown flaws and develop exploitation methods across hardened systems with limited human direction.
- The declaration materially advances OpenAI's August assessment, when it said critical capability could not yet be ruled out.
OpenAI has concluded that its upcoming Astra model meets the Critical cybersecurity capability threshold in the company’s Preparedness Framework, marking a material escalation from the precautionary assessment it published in August.
OpenAI said additional evaluations now give it enough evidence to classify Astra at the threshold, which covers models capable of developing zero-day exploits against hardened real-world systems or executing novel end-to-end attack strategies with little human direction.
The company says Astra can, when given appropriate tools and access, identify previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.
The conclusion moves beyond OpenAI’s position on 7 August, when preliminary evaluations led it to say that critical cyber capability could not be ruled out. Cyber Insider covered the additional controls introduced around Astra at that stage. The latest assessment replaces uncertainty about the classification with an affirmative finding.
OpenAI’s Critical threshold is defined in its own risk framework rather than by an external regulator. The company says a model meets it if it can identify and develop functional zero-day exploits across many hardened critical systems without human intervention, or devise and execute novel end-to-end attack strategies from a high-level objective.
Its Astra evaluation combined automated benchmarks with expert-driven testing. OpenAI reported a perfect score on ExploitBench, which tests exploit development against known vulnerabilities, and also developed an internal benchmark using more recently disclosed high-severity flaws to reduce the risk that the model had encountered the answers during training.
The classification does not mean unrestricted users will automatically receive every capability demonstrated during evaluation. OpenAI has been building additional access controls, monitoring, identity requirements, and review mechanisms around its more capable cyber models, including restrictions on higher-risk use.
The governance challenge is that defensive and offensive utility increasingly overlap. A model capable of identifying unfamiliar vulnerabilities, reasoning across large codebases, and building working proof-of-concept exploits could accelerate remediation and security research. The same capability can shorten parts of an attacker’s vulnerability-development cycle.
That creates a control problem different from ordinary content moderation. The risk is not confined to an obviously malicious prompt; it depends on what systems the model can access, the tools it can invoke, whether actions can be executed autonomously, and how reliably the provider can distinguish legitimate research from hostile preparation.
The Critical designation also raises accountability questions for organisations integrating frontier models into internal security and development workflows. Access to repositories, testing environments, credentials, package managers, and network tools can turn a capable model from an analytical assistant into an operational actor.
OpenAI’s assessment is ultimately a company statement about its own model and its own framework. Independent evaluation will remain important as capabilities move into areas where benchmark performance can translate into direct security consequences.
Even with that caveat, the change between August and September is substantive. OpenAI is no longer warning that a future model may cross its highest published cyber threshold; it is saying that one has done so. The resulting safeguards will now be tested not only by model performance, but by whether access, monitoring, and containment remain effective when the underlying capability is already present.





