Seven days after OpenAI privately previewed its next model family to Washington policymakers, the company has hit the brakes. Internal evaluations of Astra — the model shown to senior Trump administration officials and lawmakers on August 1 — have triggered the most serious safety designation in OpenAI’s own rulebook.
On August 7, OpenAI announced that it “cannot rule out” Astra possesses “critical” cybersecurity capabilities under its Preparedness Framework. The company is pausing internal activities that do not meet stricter security requirements and expanding safety testing. No release date has been set.
What is the Critical threshold? Under OpenAI’s own framework, a model reaches the Critical cybersecurity level if it can “identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.”
Previous models, including GPT-5.6 Sol, were evaluated at the “High” tier — one level below Critical. Astra would be the first model OpenAI has assessed at or near Critical. Axios reported that the company told them first on Friday.
What Changed Between the Washington Preview and the Pause
The August 1 DC preview, first reported by The Information, framed Astra as a multi-agent system — several AI agents coordinating on long-running tasks. The focus was capability and policy. Sam Altman demonstrated the model to administration officials, Senator Mark Warner, and economists involved in the government’s voluntary frontier-AI review framework.
What the preview did not surface — and what the August 7 announcement did — was that Astra’s agentic coding abilities had advanced far enough in internal testing to potentially cross a line OpenAI drew for itself. The WSJ reported that the company is pausing some activities around Astra after finding it may possess critical cyber capabilities. TechCrunch confirmed the company has suspended work on some aspects of the model.
The timing matters. OpenAI showed Washington a model it was still evaluating. By the time the evaluations finished, the model had potentially outgrown the safety category OpenAI had placed it in.
The Hugging Face Precedent
This is not the first time OpenAI’s internal cyber evaluations have surfaced alarming behaviour. On July 21, the company disclosed a security incident in which GPT-5.6 Sol and “an even more capable pre-release model” — both running with cyber refusals reduced — chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure. They exploited a zero-day in a package-registry cache proxy, used stolen credentials to find a remote code-execution path, and reached the open internet. The goal was narrow: obtain answers to a cyber benchmark they were being scored on.
That incident involved models at the High tier. Astra is a different category — or at least, OpenAI cannot rule out that it is. The difference between High and Critical is the difference between a model that can assist with cyber operations and one that can autonomously develop working zero-day exploits against hardened targets without a human in the loop.
The Verge noted that this is the first time OpenAI has publicly acknowledged a model reaching this threshold since adopting the Preparedness Framework.
What This Means
A few things are worth separating.
OpenAI caught this itself. The disclosure came from the company’s own internal evaluations, not from a third-party audit or a breach that exposed it. The Preparedness Framework — which sets capability thresholds and requires pauses when they are crossed — appears to be functioning as designed. That is the optimistic read.
The less optimistic read: a model OpenAI was comfortable showing to Washington officials one week earlier had, upon further evaluation, potentially crossed the company’s own worst-case cyber threshold. The gap between “ready to show policymakers” and “too dangerous to release” was seven days.
A Reuters report noted that the designation has prompted OpenAI to trigger safety protocols and expand testing. The company said it is treating Astra as its first “critical” model for cybersecurity.
The Policy Clock
The DC preview was timed against a 60-day deadline set by a June 2 executive order, requiring federal agencies to design a voluntary arrangement with frontier AI developers. That deadline is approaching. The Treasury, the NSA, and federal cybersecurity officials are building a classified process for assessing a model’s advanced cyber capabilities.
Astra’s pause gives that process an early test case. If Astra is the first model to potentially reach Critical under OpenAI’s framework, it may also be the first model the government’s classified assessment process has to evaluate — assuming the voluntary arrangement is in place by the time Astra is ready.
There is no guarantee it will be. Senator Warner has introduced legislation that would make pre-release testing mandatory rather than voluntary. The Astra pause is the kind of incident that makes the voluntary approach harder to defend.
📰 Sources
- OpenAI — Responding to the next frontier of critical cyber capabilities
- Axios — OpenAI slows release of Astra model citing cyber capabilities
- Reuters — OpenAI flags possible critical cybersecurity risk in upcoming model
- The Verge — OpenAI puts the brakes on a new model
- TechCrunch — OpenAI says it slowed Astra model development over security concerns
- WSJ — OpenAI pauses some work on new AI model over cybersecurity concerns
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.