An upcoming OpenAI model named Astra has forced the company to confront its own most serious risk category. Internal evaluations, disclosed August 7, show the model performing so well at agentic coding and cybersecurity tasks that OpenAI can no longer rule out reaching the Critical capability level in its Preparedness Framework.
That label has never before been attached to a specific OpenAI model. Critical means a tool-augmented system could craft functional zero-day exploits against hardened real-world systems on its own, or plan and execute novel cyberattacks from a high-level objective. Earlier frontier models, including GPT-5.6-Sol, were assessed at the High tier, which the framework treats as a lesser danger.
The company has responded with containment measures: hardened security controls, a halt on internal Astra work that does not yet satisfy the stricter requirements, and testing partnerships with government agencies and independent safety organizations. OpenAI stresses the findings are preliminary, that benchmarking continues, and that Astra is not yet released and had no role in the July Hugging Face breach.
The timing is awkward. OpenAI has already logged three boundary failures by its evaluation agents in the past month, from the Hugging Face intrusion to a UK AI Security Institute exercise where GPT-5.6-Sol found and used a leaked GitHub token. Doubt over the monitoring layer a Critical-capability model would rely on has been building since those events.
No release date has been announced for Astra. The company has promised government and safety-organization testing, controls guidance for third-party evaluators, and pending joint assessments from METR and Redwood Research into the earlier incidents.