OpenAI says its upcoming Astra model has reached a “Critical” cybersecurity capability threshold under the company’s Preparedness Framework — the first OpenAI model to be formally designated at that level. The milestone gives us an unusually clear look at how powerful future AI systems are becoming, and why some of their most advanced abilities may not be available to every user from day one.
OpenAI confirmed the assessment on 1 September after carrying out additional automated and expert-led evaluations of Astra. According to the company, the model can identify previously unknown security flaws and develop ways to exploit them across hardened systems when given the right tools and access, without requiring a person to guide every individual step.
That does not mean Astra is being released as an unrestricted hacking tool. In fact, OpenAI says the opposite: reaching the Critical threshold has triggered stronger safeguards during development and ahead of release, while access to Astra’s most advanced cybersecurity capabilities will initially be limited.
What does OpenAI mean by ‘Critical’?
Under OpenAI’s Preparedness Framework assessment for Astra, a model can reach the Critical cybersecurity threshold if it is capable of either finding and developing functional zero-day exploits across many hardened real-world systems without human intervention, or devising and carrying out end-to-end novel attack strategies against hardened targets from only a high-level objective.
OpenAI says Astra met that bar after a combination of public benchmarks, private evaluations and expert testing.
One of the headline results came from ExploitBench, where Astra achieved a 100% score on a benchmark designed to test exploit development from known vulnerabilities. OpenAI then used a newer internal benchmark containing 20 high-severity vulnerabilities disclosed between June and August 2026. On that test, Astra achieved substantially higher arbitrary code-execution rates than GPT-5.6 Sol while using fewer output tokens.
More significantly, OpenAI says Astra discovered and used two previously unknown vulnerabilities as part of an exploit chain during evaluation. The company says it is in the process of disclosing those vulnerabilities to the relevant maintainers.
OpenAI says expert-led testing went further. Astra found unknown vulnerabilities in a hardened browser and operating system, built exploit chains and demonstrated that it could move from finding a weakness to using it in a practical sequence of actions.
Why OpenAI slowed Astra’s development
The capability milestone also helps explain why OpenAI has been deliberately slowing parts of its frontier-model development over the past few weeks.
Following the separate OpenAI-Hugging Face security incident, the company paused certain frontier training, including some Astra training, for around two weeks while it hardened research environments, strengthened isolation and network controls, expanded monitoring and raised its alignment requirements.
OpenAI has been clear that Astra was not involved in the Hugging Face incident. The company has instead used lessons from that event when tightening the controls surrounding Astra and other highly capable systems.
OpenAI’s earlier development update said the company had paused or restricted workloads that did not yet meet its tougher security bar. Its latest Astra assessment says the large frontier reinforcement-learning run that had been held back was restarted on 28 August after new safety and security requirements were put in place, while some smaller experimental runs remain paused.
What safeguards is OpenAI adding?
OpenAI says Astra needs protection against two different types of risk: malicious users attempting to misuse the model, and the possibility of the model itself taking unauthorised or misaligned actions when operating with powerful tools.
The company says its safeguards include stronger refusal training for harmful cyber requests, system-level classifiers, monitoring for risky or unauthorised behaviour, stricter controls around model access and additional containment measures during training and evaluation.
That monitoring will also have consequences for legitimate users. OpenAI says some safe tasks may occasionally be slowed, paused or stopped if its systems detect behaviour that resembles cyber misuse or unauthorised activity. In ChatGPT or Codex, users may sometimes be asked to review an action before continuing; on some API surfaces, the task may stop instead.
Will Astra be the next ChatGPT model?
OpenAI has not publicly confirmed that Astra will become the next default ChatGPT model, so it would be premature to treat the two as the same thing.
What OpenAI has said is that Astra is an upcoming model and that it plans to make it available soon. The company has not announced a firm public release date.
Its most advanced cybersecurity capabilities will not simply be opened to everyone at launch. OpenAI says advanced cyber work will initially be available to a group of testers, with access through its Daybreak Blue programme following later to broaden defensive use.
That distinction is important. Future consumer AI systems may increasingly arrive with different levels of capability depending on the account, product, task and safety controls involved. A model can therefore become dramatically more capable without every user receiving unrestricted access to every part of that capability.
Why this matters beyond cybersecurity
Astra’s cyber performance is specialised, but the wider story is about AI models becoming more autonomous, more capable with tools and better able to complete long sequences of work with less supervision.
That direction is already visible across consumer technology. We have seen AI pushed further into connected-car software in our look at the latest Apple CarPlay changes, while the new Sonos 27 platform can connect with external AI assistants including ChatGPT. Astra shows what happens when the underlying models themselves make another substantial jump in capability.
For ordinary users, the most immediate change may therefore be less about a single headline feature and more about how AI services manage access. More powerful models could handle increasingly complex tasks, but we should also expect approval steps, monitoring and capability restrictions around higher-risk functions.
Reuters also reported that OpenAI had resumed its largest paused training run while continuing to hold back some smaller experiments, reinforcing that the company is now attempting to balance model progress with tighter development controls rather than simply pushing ahead at maximum speed.
Tech Torque Verdict
Astra’s Critical cyber designation is important because it moves the discussion beyond speculation about what a future AI model might be able to do. OpenAI is saying publicly that one of its upcoming systems can already perform cybersecurity work at a level that demands materially stronger safeguards.
The reassuring part is that those safeguards are being treated as a release requirement rather than an afterthought. The more challenging part is what the milestone tells us about the pace of progress. As AI systems become better at finding vulnerabilities, operating tools and completing complex tasks with less human direction, the question is no longer just how capable the next model will be. It is increasingly about which capabilities can be made broadly useful without making them broadly dangerous.




