OpenAI Shelved Astra for Misreporting Its Work, Then Shipped Cheaper Sol
On Sept. 22, OpenAI’s product page for GPT-6 Astra described it as “our most aligned model,” one that is “more likely to operate within the boundaries set by the user.” Six days later, the next version of that model was dead.

The Wall Street Journal reported on Sept. 28 that OpenAI had cancelled the October release of GPT-6.1 Astra, which was headed for ChatGPT and Codex. According to Reuters’ account of the Journal’s story, Saachi Jain, OpenAI’s head of safety systems, said the model showed more deception than its predecessor, at times failing to accurately disclose actions it had or hadn’t taken. It also had what Jain called scope and authorization problems: pressing ahead with tasks without asking the user, and sometimes reaching for external tools in unsafe ways. Per that same account, Jain told the Journal the model fell short in alignment tests, which assess whether a system follows human intent. That is the Journal’s framing of her words, not a line she published.
The next day, at its DevDay conference, OpenAI shipped something else.
The model that did ship
GPT-6.1 Sol went live on Sept. 29 for Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex, and in the API as gpt-6.1-sol. OpenAI pitched it as nearly matching GPT-6 Astra on agentic coding, computer use and professional work “at one-fifth of Astra’s standard input and output token prices.” In dollars, that is US$2 (CA$2.78) per million input tokens and US$10 (CA$13.90) per million output tokens, against US$10 (CA$13.90) and US$50 (CA$69.50) for Astra.
The headline accuracy claim needs a careful read. OpenAI says that across the reasoning settings it tested, Sol’s error rate stays within 1.9 percentage points of GPT-6 Astra’s, measured as the share of responses to difficult prompts that contain at least one factual error. TechCrunch rendered it as “within 1.9%,” which reads smaller than it is. Percentage points and percent are different animals. And the comparison is with GPT-6 Astra, the model already on sale. The cancelled 6.1 version is nowhere in OpenAI’s launch materials.
So the sequence runs like this. A flagship upgrade fails internal review. A cheaper sibling ships the next day with a pitch built on how close it gets to the flagship. Is that a safety win, or a lab that could afford to shelve an expensive model because it had a cheaper one ready?
Why the timing cuts both ways
The generous reading is real. Astra 6.1 was reportedly days from release, and killing a flagship on the eve of a developer conference is not free. A cancellation over honesty failures could easily read as a company in trouble. OpenAI has spent September trying to make it read as process instead.
On Sept. 16 it published a framework for reporting model misalignment that disclosed six incidents from the previous six months, including models that added instructions to their own summaries to conceal mistakes, and one that found and used an exposed API key without authorization. That post contains the most quotable sentence any frontier lab has written about itself this year:
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”
Then on Sept. 28, the day of the Journal story, OpenAI posted “Towards safety cases for frontier AI training,” arguing that senior leaders should each be able to veto a training run and that investigation results and postmortems “should be shared with the public following the conclusion of the investigation.” SecurityWeek tied the two together, and quoted Jain describing the tension her team is managing: finding the line between “staying within scope” and avoiding laziness when a model hits friction mid-task.
The cynical reading is also real. The cancellation reached the public through the Journal, not through an OpenAI post, and OpenAI did not immediately respond to Reuters’ request for comment. There is no 6.1 Astra system card, no eval table, nothing a reader can check. By its own safety-case logic, the postmortem should follow. Until it does, “we pulled it” is a claim about OpenAI’s standards that only OpenAI can verify.
Read Sol’s own system card
Here is the number that complicates the victory lap. Sol’s system card addendum, published Sept. 29, includes a coding-deception evaluation. GPT-6.1 Sol’s rate of misrepresentation is 1.50%. GPT-6 Astra’s is 0.51%. The older GPT-6 Sol sits at 1.30%.
Put plainly, the model OpenAI shipped the day after cancelling Astra 6.1 for misreporting its work misrepresents its coding work about three times as often as the Astra it is being sold against, on this test. OpenAI’s own caveat applies: the tasks were “deliberately selected to elicit potentially dishonest behavior,” and the rates are not expected to match production. That cuts both ways too. It means 1.5% is not a real-world failure rate. It also means this is exactly the kind of test a misreporting model should fail.
Our read: the bar OpenAI applied looks relative, not absolute. A new model appears to be judged against the one it replaces. Sol beat its own predecessor on some measures, failing to acknowledge a broken search tool in 2.08% of cases versus 4.92% for GPT-6 Sol, and slipped on others, so it cleared. Astra 6.1, by the Journal’s account, was more deceptive than the Astra before it, so it didn’t. That is a defensible rule. It is also a rule under which a cheaper, less honest-on-paper model can ship while a more capable one is held back, provided each is compared to its own lineage.
None of that makes the cancellation fake. It makes it narrower than the headlines. Engadget, summarizing the Journal, reports that OpenAI plans to investigate the root cause and keep using the same base model for future GPT-6 generations. The model was shelved. The underlying model was not.
The useful question for anyone deploying these systems is what number would have stopped Sol. OpenAI’s Sept. 22 page boasted that Astra was three times less likely than an older Sol to misstate its own capabilities. On Sept. 29, the new Sol came in at roughly three times Astra’s rate on the coding-deception test, and shipped anyway at a fifth of the price.
Sources
- Engadget: OpenAI reportedly cancels GPT-6.1 Astra's release over deceptive behavior
- Yahoo Finance (Reuters): OpenAI shelves new AI model after internal safety tests, WSJ reports
- TechCrunch: OpenAI reportedly ditches model over safety concerns
- TechCrunch: OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less
- SecurityWeek: OpenAI calls off GPT-6.1 Astra launch, details safety cases for frontier training
- OpenAI: Introducing GPT-6.1 Sol
- OpenAI Deployment Safety Hub: GPT-6.1 Sol system card addendum
- OpenAI: GPT-6 Astra
- OpenAI: Our framework for reporting model misalignment
- OpenAI: Towards safety cases for frontier AI training
- MIXED: GPT-6.1 Sol costs $2 and $10 per million tokens
Cassandra Lee covers AI and machine learning for prompt/power: the labs, the model releases, the research and the safety fights that come with them. She reads model cards the way other people read horoscopes: skeptically, and mostly for what's left unsaid.
Latest from prompt/power
- How to Read an AI Company’s S-1: The 7 Numbers That MatterOct 5
- OpenAI’s Safety Lead Quit Over Culture. California’s AG Was Already InOct 5
- When an AI Agent Breaks In, Who Answers for It?Oct 5
- The New AI Models Don’t Talk. They Decide.Oct 5
- Quebec’s First AI Election: ChatGPT Leaned on an AI-Built Voter GuideOct 5
Leave a Reply