GPT-5.6 shipped with a government green light. Its cyber guardrails lasted hours.
It took the United States government roughly two weeks of supervised access to decide GPT-5.6 was safe enough to release. By one published account, it took the United Kingdom’s red team about six hours to break it.
Lay the dates end to end and the shape of the problem is hard to miss. On July 8, CNBC reported that OpenAI would publicly release the GPT-5.6 family, ending the government-requested limits that had confined the models to vetted partners. On July 9, OpenAI shipped Sol, Terra, and Luna to ChatGPT, Codex, and the API, calling GPT-5.6 its “strongest cybersecurity model yet.” Within hours of public access, the UK AI Security Institute’s red team lead, Xander Davies, posted findings of universal jailbreaks in the cyber domain. The same day, OpenAI doubled the top prize in its universal-jailbreak bounty program to $50,000.
Approval, launch, break, bounty. Four days.
What Washington signed off on
The release limits were real, and unprecedented. On June 26, OpenAI said it would restrict the GPT-5.6 preview to “a small group of trusted partners” at the request of the Trump administration, under an executive order asking frontier labs to voluntarily give the government up to 30 days of pre-release access to models with advanced cyber capabilities. OpenAI complied and grumbled in the same breath: “We don’t believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them.”
The review ran through the Commerce Department’s Center for AI Standards and Innovation, with Commerce Secretary Howard Lutnick, Treasury Secretary Scott Bessent, and National Cyber Director Sean Cairncross involved, according to The Next Web. OpenAI described the process as a “collaborative back and forth”; Sam Altman said the company made “many changes” and sent technical staff to Washington. It was the first US frontier model launched under a government-managed access framework, and by July 8 the framework said yes.
Note what that yes was. A cooperative, pre-deployment review conducted with the model’s maker in the room, weighing capability against the safeguards OpenAI presented. Not an adversarial attempt to tear those safeguards off.
What the system card promised
OpenAI’s own paperwork is specific, which makes it checkable. The GPT-5.6 system card classifies all three models as High capability in Cybersecurity under the company’s Preparedness Framework (pp. 44–45), and says the accompanying safeguards are “our most robust yet,” blocking “roughly ten times more potentially harmful activity” than prior models (p. 1). The capability numbers underneath are striking: Sol saturates capture-the-flag evaluations at 96.7% (p. 47), though it “did not independently produce a functional full chain exploit” against real-world targets in critical-threshold testing (p. 50).
The jailbreak-robustness section is quieter. Sol “performs comparably to recent predecessors” on worst-case defender success rate (p. 11). That is not a boast. It is an actuarial note: the strongest cyber model OpenAI has ever shipped is about as resistant to jailbreaks as the models before it. The safeguards got stronger; the lock on them did not.
The card also points security professionals who want Sol with fewer cyber restrictions toward a formal channel, the Trusted Access for Cyber program (p. 45). Hold that thought.
“Universal” is the load-bearing word
What AISI found, per Fortune’s July 10 report and the institute’s own publication the next day, were universal jailbreaks in the cyber domain: prompt-level breaks that transfer across sessions and tasks, unlocking “long-form agentic task completion in domains like vulnerability discovery and exploit development.” Not a one-off trick that coaxes a single bad answer. A reusable key. AISI researchers had privileged access OpenAI users don’t get, and the institute said such jailbreaks were “often developed” within hours; TechRepublic reports one universal jailbreak took six hours to build. Davies was direct about the access asymmetry: the jailbreaks “are still findable without this access, just slower.”
Why the distinction matters: a one-off exploit dies when it’s patched. A universal jailbreak defines a class, and patching instances doesn’t kill the class. Security researcher Stanislav Fort, quoted in Technobezz’s coverage, put it plainly:
Patching “only closes those specific attack instances, not the category as a whole.”
AISI itself said it “expects further red teaming to surface similar jailbreaks.” OpenAI’s response, per Fortune: there is “no such thing as perfect security,” safeguards are layered, remediation is rapid. The company says it “worked to reproduce and mitigate the specific jailbreaks reported by UK AISI.” The specific jailbreaks. Fort’s point stands unanswered.
And this failure class has a recent, severe precedent, which is Fortune’s framing and worth attributing as such: in June, after Amazon researchers found a jailbreak in Anthropic’s Fable 5, the US government imposed export controls that forced Anthropic to disable the model entirely on June 12. The controls were lifted July 1, after two weeks of negotiation. Same category of flaw, per Fortune. In one case it took a frontier model offline; in the other, so far, it has produced a statement about layered security.
Two governments, two tests
Here is the analysis, flagged as ours: nothing in this sequence requires anyone to have lied. The Commerce review and the AISI red team were administering different exams. “Approved for release” measured whether the model plus its safeguard architecture, presented cooperatively over weeks, cleared a deployment bar. “Safe at the jailbreak layer” is an adversarial question, answered after launch, by a different government, on a different clock. Both verdicts can be simultaneously accurate. That’s the uncomfortable part. The gap between the two tests is now the operative AI policy story, because the US built a front door for pre-release review at the exact moment the UK demonstrated that the back door opens in an afternoon.
The capability side of the ledger is not hypothetical, whatever you think of benchmarks. The usual caveats apply to OpenAI’s own numbers (an 80 on the Artificial Analysis Coding Agent Index, “2.8 points above Fable 5,” is a vendor-chosen stat on a vendor-chosen index). But at the AtCoder World Tour Finals in Tokyo this month, an OpenAI reasoning model that engineer Borys Minaiev described as “comparable to GPT-5.6” swept all five algorithm problems for 8,300 points; the best of fourteen elite humans scored 4,300, and no human solved problems C or E at all. “Humanity has not prevailed,” wrote Psyho, the programmer who beat OpenAI’s model at the same event last year. Contest wins aren’t exploit chains. But a week before the launch, Sysdig documented JadePuffer, the first known ransomware operation in which an LLM agent autonomously ran the attack, encrypting 1,342 configuration items and self-correcting a failed step in 31 seconds. Agentic cyber capability in the wild is a current event.
Which brings us back to the bounty. The $50,000 prize OpenAI doubled on launch day is, per TechRepublic, its Bio Bounty: a private, NDA-gated program for universal jailbreaks of the biological safeguards, application required. The cyber domain, where AISI just demonstrated the break, has its own channel in the system card, the one for professionals who want Sol with fewer safeguards through official enrollment.
For a few hours on July 9, according to the UK government, the enrollment requirement was a prompt.
Cassandra writes about technology as a cultural force — what it does to how we live, work, and understand ourselves. She has a background in cognitive science and too many browser tabs open. Based in Vancouver.
Leave a Reply