OpenAI’s Astra Is the First Model It Calls Too Dangerous — On Purpose
OpenAI released GPT-6 Astra on September 3, 2026, and the launch followed the familiar script: a flagship model billed, in the company’s own words, as the world’s most intelligent and aligned model, state of the art on computer use, browsing, software engineering, cybersecurity, science, and professional work. If you have watched a model launch before, you can recite the genre. The interesting part of this one is not the superlatives. It is that OpenAI, for the first time, rated one of its own products as hitting the Critical threshold for cybersecurity under its Preparedness Framework — and shipped it with the sharpest edges filed down.
What’s real. The Critical classification is not marketing garnish; it is a specific, self-imposed designation. Under OpenAI’s framework, that threshold is reached when a model can, roughly, find and build working zero-day exploits against hardened real-world systems without a human in the loop, or run novel end-to-end attacks against hardened targets from only a high-level goal. According to OpenAI’s disclosures as reported by InfoQ, Astra in testing discovered previously unknown vulnerabilities in a browser and an OS kernel. It built a working exploit achieving unsandboxed code execution in about 29 hours against a browser build lacking production mitigations, then took roughly 12 more hours to adapt it to the official stable release. The Hacker News reports it scored 100% on ExploitBench.
That is the load-bearing news, and it has been widely reported: OpenAI says its own model crossed a line it had drawn in advance, and so it restricted the released version to defensive tasks — secure code review and patching — while refusing to generate proof-of-concept exploits. According to InfoQ’s read of OpenAI’s system card, the safeguards include stricter isolation, encrypted model checkpoints and monitoring of the model’s full reasoning trajectories. Access to the model’s cyber capabilities runs through a gated program OpenAI calls Daybreak. Bloomberg and multiple security outlets led with the added cyber guardrails rather than the benchmarks, which is the bar that matters.
What’s marketing. Almost everything with a number next to it. OpenAI reports Astra at 97.6% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, 72.6% on OSWorld 2.0, and 59.3% on Agents’ Last Exam. Those are the company’s own figures, on benchmarks that OpenAI selects and, in some cases, whose test conditions it controls. They are not independently reproduced results, and readers should treat them as vendor claims until third parties run the model themselves. A perfect 100% on any benchmark should raise an eyebrow, not settle an argument — it usually means the benchmark is saturated, not that the problem is solved. The recurring “world’s best” and “most intelligent and aligned” language is positioning, and “aligned” is doing an enormous amount of unearned work in a sentence that is also announcing the model is dangerous enough to withhold features.

The pricing is real and worth noting. Astra runs about $10 per million input tokens and $50 per million output tokens at the standard tier, with a faster mode at a premium and higher rates for long-context calls, per OpenAI and Microsoft Foundry listings. Frontier capability is not getting cheaper at the top. Availability spans ChatGPT’s paid tiers, the OpenAI API as gpt-6-astra, Microsoft Azure, and AWS Bedrock.
Why it matters. For years the critique of AI safety messaging was that it was theater — dire warnings from the same companies racing to ship. Astra complicates that. Here a lab published a capability threshold, then said its product met it, then constrained the product accordingly. That is closer to the model of “measure the risk, then act on the measurement” that safety researchers have asked for. It is also, conveniently, a story that makes the model sound extraordinarily powerful. Both readings are available, and both are probably correct: this is a real safety decision and an excellent piece of marketing, and the tech press should be able to hold those two facts at once.
The test is what happens next, and OpenAI has already said what it plans: “Through OpenAI Daybreak, we plan to expand access and roll out less restrictive safeguards in the coming weeks.” The precedent is also spreading inside OpenAI: GPT-6.1 Sol, launched at DevDay on September 29, is likewise rated Critical and inherits Astra’s safeguards. The question is no longer whether the limits loosen, but for whom and under what verification. If Daybreak becomes a meaningful gate, the Critical rating will have been real policy. If it becomes a paywall, it will look like a launch-week flourish.
Sources
- OpenAI, "GPT-6 Astra: A new generation of intelligence"
- InfoQ, "GPT-6 Astra is the First Model OpenAI Classifies as Critical for Cybersecurity"
- The Hacker News, "GPT-6 Astra Scores 100% on ExploitBench as OpenAI Blocks PoC Exploit Requests"
- NBC News, "OpenAI debuts GPT-6 Astra, says it triggered security measures"
- CNBC, "OpenAI announces rollout of GPT-6 Astra model" (Sept 3, 2026)
- Bloomberg, "OpenAI rolls out GPT-6 Astra model with added cyber guardrails" (Sep 3, 2026)
Cassandra Lee covers AI and machine learning for prompt/power: the labs, the model releases, the research and the safety fights that come with them. She reads model cards the way other people read horoscopes: skeptically, and mostly for what's left unsaid.
Latest from prompt/power
- How to Read an AI Company’s S-1: The 7 Numbers That MatterOct 5
- OpenAI’s Safety Lead Quit Over Culture. California’s AG Was Already InOct 5
- When an AI Agent Breaks In, Who Answers for It?Oct 5
- The New AI Models Don’t Talk. They Decide.Oct 5
- Quebec’s First AI Election: ChatGPT Leaned on an AI-Built Voter GuideOct 5
Leave a Reply