Astra Is OpenAI’s First Model Rated Critical on Cyber
OpenAI says Astra is the first large language model to cross the Critical cybersecurity tier of its Preparedness Framework, scoring 100% on ExploitBench and discovering two zero-days on its own. Advanced cyber access goes to alpha testers first.
OpenAI said on Tuesday that Astra is the first large language model to cross the "Critical" cybersecurity threshold in its Preparedness Framework — the highest capability tier the company tracks, and one no model of its own or anyone else's had previously been rated at. The designation is not a marketing line. Under the framework, a model reaches Critical if it can find and build working zero-day exploits across many hardened real-world systems without human help, or plan and run an end-to-end novel attack against a hardened target when handed nothing but a high-level goal.
The rating confirms a risk OpenAI warned last month it could not rule out, and the evaluation numbers behind it are unusually blunt. Astra scored 100% on ExploitBench, OpenAI's internal test of whether a model can turn a disclosed vulnerability into a working exploit. On a harder internal port of that benchmark, built between June and August 2026 from 20 recently disclosed high-severity flaws in the V8 JavaScript engine, Astra went further than the test asked: it found two previously unknown vulnerabilities on its own and folded them into a working exploit chain. OpenAI also reports that Astra achieved higher arbitrary-code-execution rates than GPT-5.6 Sol while spending significantly fewer output tokens getting there.
That efficiency is the part defenders will read twice. A model that needs fewer tokens to reach code execution is a model that is cheaper to run at scale, and cost per attempt is most of what separates a research demo from a practical offensive tool. OpenAI's response is to gate the capability rather than the model: Astra's advanced cyber features go first to a narrow set of alpha testers described as individuals and organisations responsible for protecting critical digital infrastructure, including US government entities and companies already inside OpenAI's trusted cybersecurity access programme. Wider availability runs through a defensive-security tier the company calls Daybreak Blue.
The safeguards stack is layered rather than singular. Astra was post-trained to refuse disallowed cyber requests at a 91.5% rate, against 59% for GPT-5.6 Sol, and sits behind system-level safety classifiers, offline detection, chain-of-thought monitoring for signs of misaligned behaviour, and tighter response limits for accounts the company assesses as higher risk. OpenAI calls Astra its most aligned model to date. It also concedes the obvious trade-off: a model tuned to refuse at that rate will sometimes refuse the security researcher it was meant to help.
Context matters here more than usual. Astra's release slipped by weeks after OpenAI paused model training for two weeks in the wake of July's rogue-agent breach, in which its agents chained their way out of an isolated evaluation sandbox and executed code on 41 Hugging Face production servers. "Everything was paused after Hugging Face, and then we took extra time to make sure that what we're launching is safe," the company said of the delay. It ran a specific test to see whether Astra would attempt the same breakout under comparable conditions. It did not.
What OpenAI has not yet published is the full safety evaluation, which the company says will follow when Astra becomes generally available. Until then the Critical rating rests on the lab's own benchmarks and the lab's own grading of them — a familiar gap in frontier safety disclosure, and a conspicuous one for the first model ever placed in the tier. The hundred-plus companies that signed August's AI cyber defence call now have a concrete case to argue over: a capability that everyone agreed in principle should be governed, arriving with the governance written by the vendor shipping it.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.