OpenAI Priced Containment at One Fifth of Its Compute Bill
Astra is the first model to cross OpenAI's own Critical cybersecurity threshold. The controls OpenAI deployed are reactive by design. Release is weeks away. Read that in order.

OpenAI's Preparedness Framework says a model reaches Critical cybersecurity threshold when it can autonomously identify and develop zero-day exploits across hardened real-world systems without human intervention. The consensus was that this line meant something. OpenAI crossed it, published a blog post, and kept building toward release.
This is not a story about a rogue AI. It is a story about what safety frameworks actually do when the cost of slowing down becomes real.
What OpenAI Built and What It Deployed to Contain It
Under OpenAI's framework, a model reaches Critical if it can identify and develop functional zero-day exploits in many hardened real-world critical systems without human intervention. OpenAI's evaluations of Astra indicated performance strong enough that it cannot rule out Critical capability. In nearly three years, no prior model had triggered the Critical threshold's development-stage requirements.
Detection is not containment.
What followed was a two-week pause on reinforcement-learning training, plus monitoring that consumes 20% of the inference compute it watches, plus a rule where three teams get paged and the run stops if nobody clears the flag inside 30 minutes.
Translation: if Astra starts mapping your network, OpenAI has half an hour to notice and someone has to be awake.
The 20% compute tax is honest. The problem is what it buys. Monitors evaluate the model's chain of thought and trigger security review. A model that discovers a zero-day will have already discovered it before the monitor sees the thought that produced it. Detection after the fact is incident response, not safety control.

The Monitorability Problem
The surveillance approach depends on reasoning staying legible. OpenAI's own research documents the opposite: increased parameter counts mean more inference compute per token, giving cognition more opportunity to hide in activations rather than chain of thought. Larger models are less monitorable at fixed capability levels.
The controls assume visible reasoning. The scaling curve makes visible reasoning less likely with every generation. OpenAI's research also notes: if strong supervision applies to chain of thought, models can learn to hide intent while continuing to misbehave.
No external independent verification of OpenAI's control claims exists.
The Clock
OpenAI has expanded internal testing of Astra under partner names like mozaik-alpha-fdm and ultima-alpha, with observers noting release might be near. Expanded partner testing under a named checkpoint is what the last three OpenAI model releases looked like two weeks before they shipped.
OpenAI says it started implementing stricter security controls for testing, including isolated environments and universal monitoring. A technical staff member said the company is "consciously slowing down research to enhance security."
Slowing down is not stopping. The checkpoint is named. Partner testing is live.
If you are running untrusted code in a shared environment, your blast radius changed. Astra does not ask for network access. It discovers what is reachable.
The real question is not whether OpenAI will release a Critical model. It already decided that. The question is why Critical designation stopped nothing, and what that tells you about every model that follows.
What to Watch
Watch for the actual release date and usage restrictions. Watch for the first published autonomous attack on a hardened testbed: the pattern suggests disclosure within 90 days of release. Watch whether CISA, the SEC, or Congress asks on record why Critical did not trigger a release halt.
For enterprise security teams, developers building on OpenAI's API, and policymakers debating voluntary AI safety governance, the answer to one question is writing itself: when voluntary safety frameworks cost something commercially, do they hold?
Your zero-day patching team is not ready for machine-speed vulnerability discovery. Neither is your incident response window.
- Responding to the next frontier of critical cyber capabilities | OpenAI
- Pacing model development in an era of cyber-critical capabilities | OpenAI
- OpenAI's Upcoming Astra Model Raises Autonomous Cyberattack Concerns | SecurityWeek
- OpenAI Pauses Astra After Tests Reveal Autonomous Zero-Day Exploit of Hardened Systems | TechTimes
- OpenAI Warns Astra AI Model May Develop Zero-Day Exploits | GBHackers
- Exclusive: OpenAI slows release of Astra model citing cyber capabilities | Axios
- OpenAI's Astra can do a researcher's week of work. That's the problem. | The New Stack
- First outputs from GPT-6 Astra model from OpenAI | TestingCatalog
- OpenAI Astra: First Outputs Leak Under 'mozaik-alpha-fdm' | OrcaRouter
- Evaluating chain-of-thought monitorability | OpenAI
- Detecting misbehavior in frontier reasoning models | OpenAI
- OpenAI Slows Frontier AI Training as Astra Nears Critical Cyber Threshold | eSecurity Planet
- OpenAI Flags Possible Critical Cybersecurity Risk in Astra AI Model
- OpenAI says Astra could reach 'critical' cyber capability ...
- OpenAI flags possible critical cybersecurity risk in upcoming model Astra | Tech News - Business Standard
- OpenAI Warns Astra AI Could Develop Zero-Day Exploits and Launch Autonomous Cyberattacks
- OpenAI's Astra Model: What We Actually Know So Far | MindStudio
- OpenAI AGI Timeline Indicates Astra Model Coming This Year - Geeky Gadgets
- OpenAI Astra: every confirmed fact about the next model | eesel AI
- Astra Release Date: What Is Actually Known | How To Use Astra
- OpenAI Astra: Quantum Mathematics and Cybersecurity Risks | by SOCFortress | Aug, 2026 | Medium
- Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
- Chain-of-Thought Monitorability in AI: OpenAI Introduces New Evaluation Framework for Transparent Reasoning | AI News Detail
- Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
- Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
- Lentils on X: "🚨 Major Scoop: OpenAI has just expanded internal testing of GPT Astra, codenamed "mozaik-alpha-fdm", signifying release might be near AND of course like previous times, we have the FIRST EVER public outputs of it for y'all😉 Both outputs are zero-shot on Max effort. Frontend" / X
- Leaked Astra Outputs Show OpenAI's Stunning Visual Advances / X
- 🚨 AI News | TestingCatalog on X: "ASTRA 🔥: First outputs of the upcoming unreleased model from OpenAI have arrived, and these results are stunning! > Internally named "mozaik-alpha-fdm", rumored to be a checkpoint of the "Astra" model, mentioned by OpenAI earlier as a substantial step-forwar… / X
- XIVIX on X: "🚨 OpenAI has just dropped its first internal Astra checkpoint: mozaik-alpha-fdm. Here are the first two outputs, both generated 1-shot on Max effort. On Max the model tends to think ALOT more than Sol, but it has a ton of attention to detail as evident by the outputs below." / X
- Wes Roth on X: "🚨 GPT ASTRA might be getting close. OpenAI has reportedly expanded internal testing of GPT Astra, codenamed: “mozaik-alpha-fdm”" / X
- Chetaslua on X: "🚨OpenAI Astra : Frontend is solved Attention to detail man , we are so back zero-shot on Max effort < mozaik-alpha-fdm > this is the first time in history now many idea guys will have edge over the execution guy as this ai can execute what ever you think in reality" / X
- OpenAI Astra Rumor Tracker: Demos, Dates and Evidence