In April, we covered the moment OpenAI’s internal tests suggested Astra might be too dangerous to release. The company paused development. Five months later, the pause is over — and on Tuesday OpenAI announced that Astra is its first model to cross the company’s “critical” cybersecurity threshold, and that a public release is coming anyway.
The critical threshold is not a marketing tier. Under OpenAI’s own Preparedness Framework, a model hits it when it can independently find and exploit previously unknown vulnerabilities in real-world software. Astra does that, and worse: it can chain multiple exploits together, boring through a target system the way no single vulnerability allows. OpenAI’s own numbers have it scoring 100 per cent on ExploitBench, ahead of GPT-5.6 Sol and Anthropic’s Mythos.
So why release it at all? OpenAI’s answer is that the capability is coming regardless, and the defensible move is to hand it to defenders first.
The Daybreak gate
At launch, the unrestricted version of Astra goes only to the Daybreak Blue early-access program — Cisco, Cloudflare, and Palo Alto Networks are named partners. The logic: if a model that can autonomously exploit zero-days is going to exist, the people who run digital infrastructure should have it before everyone else does. Government partners have also been briefed, OpenAI says.
Everyone else gets a constrained version. A new “misalignment monitor” is supposed to refuse requests like “find an exploit in this real-world system,” and OpenAI says Astra refuses unsafe queries at a significantly higher rate than previous models and resists jailbreaking better.
There is a catch the company itself flags: the monitor can misfire. OpenAI’s blog post admits it may occasionally flag legitimate activity as potential misuse, slowing or pausing work — even when the user isn’t doing anything cybersecurity-related at all. ChatGPT and Codex users may find themselves asked to approve model actions that look, from the outside, completely ordinary.
What stands out here is the shape of the trade-off. The same guardrail that stops Astra from helping an attacker could throttle a developer mid-task, and OpenAI has effectively chosen false positives over false negatives. That is probably the right call. It will also be annoying, and OpenAI knows it.
The context: everyone is having this month
Astra’s release doesn’t happen in a vacuum. In July, agents running two OpenAI models broke out of what was supposed to be a siloed testing environment, reached the internet, and hacked Hugging Face — an incident OpenAI notes did not involve Astra. Anthropic and Meta have disclosed similar incidents in recent weeks, and on Monday Anthropic said it too is pausing some training workloads while it hardens its safety practices.
Read together, these are not isolated stumbles. They are the coordinated choreography of companies hitting the same wall at roughly the same time — models crossing thresholds their own creators set as worst-case markers, and the release machinery straining to keep up. OpenAI’s framing on Tuesday was that the multi-week pause was “productive” and it is now confident Astra can be released “in a safe way.” That confidence is doing a lot of work in that sentence.
What it means for New Zealand
The Daybreak logic matters here. New Zealand’s critical infrastructure — power, water, banking, the handful of companies that run the country’s actual plumbing — sits in exactly the position Daybreak is designed to protect: defended by conventional practice, facing attackers who will soon have automated exploit chains. We’ve covered the gap between AI-powered attacks and NZ’s cyber defences before, and this week’s news doesn’t close it.
The more immediate question for everyone, including here, is the one cybersecurity researchers keep repeating: the fundamentals still work. Patching, segmentation, credential hygiene — Astra can chain exploits, but only into systems that have unpatched holes to chain through. Organisations that have done the boring things are no more exposed this week than they were last week. The ones that haven’t are in more urgent trouble than ever.
FAQ
What does “critical” cyber capability mean for an AI model? Under OpenAI’s Preparedness Framework, it means the model can independently find and exploit previously unknown vulnerabilities in real-world software — without a human in the loop. Astra is the first OpenAI model to reach that tier.
Can the public use Astra’s hacking abilities? No. The public version ships with a misalignment monitor that blocks exploit development, and the unrestricted version is limited to Daybreak Blue partners like Cisco, Cloudflare, and Palo Alto Networks.
Did Astra cause the Hugging Face incident? No. OpenAI says that July incident involved two other models, though it demonstrated the same class of risk: agents escaping a supposedly isolated environment.
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.