OpenAI Can't Rule Out Critical Cyber Risk in Astra
OpenAI slowed work on its next frontier model after internal evaluations suggested Astra may cross the Critical cybersecurity threshold — a level no model has triggered before. Weights are locked down, some internal work is paused, and outside evaluators are being called in.
OpenAI said on Friday that it has slowed development of Astra, its next frontier model, after internal evaluations suggested the system may reach the Critical cybersecurity capability level defined in the company’s own Preparedness Framework — a threshold no model has triggered since the framework was written in 2023. The company was careful about its wording: it has not declared Astra Critical, only that it “cannot rule out” that classification while benchmarking continues.
The distinction matters, because the Critical bar is deliberately extreme. Under the framework, a model qualifies if it can identify and develop working zero-day exploits across severity levels in hardened, real-world systems without human intervention, or devise and execute novel end-to-end attack campaigns against hardened targets when handed nothing but a high-level goal. OpenAI describes that as a qualitatively new threat vector with no ready precedent, and unlike most of its risk tiers, Critical triggers safeguards during development rather than only at deployment.
Those safeguards are now in force. According to the disclosure, reported by TechCrunch and summarised by Unite.AI, OpenAI moved Astra into isolated testing environments with restricted network and tool access, added encryption around model weights, paused internal Astra work that did not meet the strengthened controls, and turned on monitoring that inspects chain-of-thought for risky action and can interrupt a run mid-task. Higher-capability checkpoints are being executed in sandboxes. The company also said it will bring in government agencies and selected AI safety organisations to run independent testing, and will hand those partners its recommended security controls.
What OpenAI did not provide is a timeline. Astra remains unreleased, no resumption date was given, and the company said the evaluations are preliminary with assessment still underway. Every previous frontier model it has shipped, including GPT-5.6-Sol, evaluated at High or below on the cybersecurity axis — so there is no in-house playbook for what happens next, only the framework’s instruction that Critical capability demands safeguards before the model goes anywhere near customers.
The disclosure lands in a season that has made the abstract risk concrete. Regulators finalised mandatory pre-deployment cyber testing for frontier models earlier this month after a run of agent-driven breaches, and Hugging Face’s forensic timeline of an autonomous intrusion showed what an agent with tool access can accomplish in a few days. Astra is the same model family OpenAI recently promoted for solving long-open mathematics problems; the agentic coding ability behind that result is precisely what makes the security evaluation hard.
Reaction has split along familiar lines — some lawmakers and researchers read a voluntary pause as evidence that self-governance can work, others as evidence that a company grading its own homework is exactly the problem. Both readings will stay untestable until the external evaluations report back. The number worth watching is not whether Astra ships, but whether OpenAI publishes the benchmark results that led it here, and in enough detail that outside reviewers can argue with them.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.