Live
Abstract illustration: a glowing gem held up on a platform of pillars, one of them cracked
AI & ML

What Is an AI ‘Safety Case’? How Labs Decide a Model Is Too Risky to Ship

In the last week of September 2026, OpenAI had a model ready to ship and decided not to. GPT-6.1 Astra was slated for ChatGPT and Codex in October. Then, as The Wall Street Journal first reported, the company scrapped the release after internal testing found the model was more deceptive than its predecessors, and that it wasn’t honest with testers about which actions it had and hadn’t taken.

The decision itself took a sentence to report. The machinery behind it is harder to see, and it is the same machinery every frontier lab now says it runs before a model reaches the public. This is a guide to that machinery: what a “safety case” is, who tests the model, who checks the testers, and where the word “independent” starts to strain.

A safety case is an argument, not a test

The term comes from industries that learned the hard way. Nuclear plants, aircraft and rail signalling systems are approved on the basis of a written argument that the risk is acceptable, with every claim tied to evidence. A team at the UK’s AI Security Institute adapted the idea for AI in a 2024 paper, defining a safety case as “a structured, evidence-based argument aimed at demonstrating why the risk associated with a safety-critical system is acceptable.”

The structure matters more than the length. A top-level claim (“this model cannot meaningfully help an attacker break into networks”) is broken into narrower sub-claims, and each sub-claim has to point at something concrete: an evaluation result, a red-team log, a description of a safeguard. The AISI paper’s worked example is what researchers call an inability argument. You don’t argue the model will behave; you argue it can’t do the dangerous thing even if it tried.

Inability arguments are the easy kind. They stop working the moment a model is capable enough to do the dangerous thing, at which point the case has to shift to arguments about control (we can contain it), monitoring (we’d catch it) or alignment (it won’t want to). Each of those is harder to prove than the last.

What OpenAI now says a case should contain

The same week Astra was pulled, OpenAI published a proposal arguing that safety cases should be required before frontier reinforcement-learning training runs proceed, not only before launch. The company calls the full version “an aspirational north star” and concedes that AI is nowhere near the rigour of aviation or nuclear power.

The proposal sorts the work into three layers. Technical safeguards cover alignment training, containment (sandboxing and hardened research infrastructure) and monitoring to catch misaligned actions early. Operational rules cover dissents and pre-mortems, sign-off by senior leaders with veto power, pause protocols, rollback and documentation of whatever risk is left over. The third layer is what happens after something goes wrong: root-cause analysis, postmortems and, per OpenAI, public disclosure of the findings.

One line in it deserves more attention than it got. OpenAI writes that auditors “should be provided with sufficient access to verify that the claims of the safety case are valid and sound, and to raise gaps if found.” Hold onto that phrase, sufficient access. Most of the fights in this field are about how much of it outsiders actually get.

Red teams, evals and the people who run them

A layered square core sits inside a dashed ring as dozens of fine probe lines converge on it from every direction; three highlighted pink probes reach past the ring into the core.
Illustration: prompt/power

The evidence inside a safety case mostly comes from two places. Evaluations are standardized tasks that measure a capability: can the model find a software vulnerability, follow a lab protocol, speed up an AI research project? Red-teaming is adversarial: people, and increasingly other models, try to make the system misbehave and record what works.

Labs do most of this in-house. Some of it goes to outside groups, the best known of which is METR, a nonprofit that runs pre-deployment evaluations of frontier models. Its published summaries are the closest thing the public gets to a third-party view, and they are unusually candid about their own limits.

Take METR’s summary of its evaluation of Anthropic’s Claude Opus 5.5, published Sept. 22, 2026. Testing ran through API access “granted over a period of 10 business days.” The work was done “under an unpaid agreement.”

“We drafted the initial summary, and then Anthropic had the opportunity to review and edit the text.”
— METR, on its Claude Opus 5.5 evaluation

METR also says it relied on “an additional source of information which we are not able to disclose at this time.”

None of that is a scandal. METR discloses it precisely so readers can weigh the findings. But it describes what third-party evaluation usually means in practice: a fixed window, access on the lab’s terms, and a report the lab sees before you do.

The White House accord, and who picks the auditor

On Sept. 29, 2026, Anthropic, Google, Meta, Nvidia, OpenAI and xAI signed a voluntary accord at the White House. CoinDesk’s reading of the one-page document lays out four layers: internal monitoring during training and deployment, an independent auditor’s assessment of safety controls, review of the audit findings by a board committee, and remediation of problems overseen by the board. IAPP reported that signatories committed to set up “an independent board-level committee to review internal progress reports on company controls.”

What the accord leaves out is the more useful list. CoinDesk found no enforcement mechanism, no penalties, no deadline and no requirement to publish audit results, and companies choose their own auditors without having to say who they are. President Trump called the commitments “morally binding,” telling reporters, per Axios, that the companies are “really going to be policing each other.”

An audit is only as independent as the process that picks the auditor, sets the scope and decides what the public gets to read. That’s the structural problem. Financial auditing has the same conflict, since companies pay their own auditors, but it is wrapped in licensing, liability, mandated disclosure and regulators who can bar a firm from practice. AI auditing in late 2026 has none of those yet.

The labs’ own policies show the range. Anthropic’s Responsible Scaling Policy, version 3.1, requires outside review only when a risk report covers models past its automated AI research threshold and is “significantly redacted.” Anthropic selects the reviewers “in consultation with the Board and LTBT” (its Long-Term Benefit Trust), screening for expertise and independence from its financial interests, and asks them for public commentary within 30 days. That’s more than nothing. It is also a system in which the reviewed party chooses the reviewer.

When the deadline is self-imposed

The other test of a safety process is whether a lab keeps the promises it set for itself. Anthropic’s public Frontier Safety Roadmap committed to a prototype “by September 30, 2026 of provable inference, a technique for reliably, provably ‘signing’ AI model outputs in a way that makes them attributable to a specific set of model weights.” The point is to detect an attacker who tampers with a model after training.

That date was already a revision. Anthropic’s roadmap updates page records that on May 5, 2026, it moved the Phase 1 target from May 15 to Sept. 30. Forkast noted on Sept. 30 that, by close of business that day, Anthropic had published nothing confirming the milestone, and described Phase 1 as a planning and inventory step covering components, costs and timelines rather than working software.

The roadmap is candid about what it is. Anthropic calls the goals “subject to change” and says it will “strive to avoid situations where we revise the goals in a less ambitious direction because we simply can’t execute.” Strive is the operative word. Nobody outside the company can compel the update, and no penalty attaches to a missed date. The only enforcement is that people check.

How to read the next launch, or the next cancellation

When a lab ships or pulls a model, four questions cut through most of the press release:

  • What kind of argument is it? “It can’t do X” is a stronger claim to verify than “we would catch it if it did X.”
  • Who tested it, for how long, with what access? Look for named outside evaluators and the terms they disclose.
  • Who saw the report before publication? If the lab edited the evaluator’s summary, that’s normal, and worth knowing.
  • Was there a date, and was it met? Self-set deadlines are the cheapest commitments to make and the easiest to quietly move.

Astra is a useful reference point because its failure was visible in exactly the place a safety case is supposed to look. The model, by the account of OpenAI head of safety systems Saachi Jain in SecurityWeek’s report, “fell short on scope and authorization, and on how it tells users what type of work it has done.” A model that misreports what it did is a model whose test results are harder to trust. The evidence a safety case rests on was, in part, the model’s own account of itself.

Sources

// Contributor, AI Explained
Catherine Crowe

Catherine Crowe covers AI explained for prompt/power: the plain-English guides that break down how the technology works, what the jargon means and what it changes for everyday people. Originally from Canada, she writes from New Zealand.

Latest from prompt/power

  1. DeepSeek V4.1 Flash Cut the US AI Lead to 3%. What LiveBench MeasuresOct 6
  2. Ben Affleck Calls AI Job Fears ‘Propaganda.’ He Sold an AI Firm to NetflixOct 6
  3. Cohere’s North 2 Puts AI Agents on a Budget. Toronto Bets on BoringOct 6
  4. How to Stop ChatGPT, Claude, Gemini and Meta AI From Training on Your ChatsOct 6
  5. Musk Is a Trillionaire Again After SpaceX Stock Jumps 7.6%Oct 6

Leave a Reply

Your email address will not be published. Required fields are marked *