Live
Abstract illustration of a tall isometric stack of thin translucent slabs; only the bottom few glow solid pink, while the rest are faint outlines and loose sheets drift off the top, on a deep purple background.
AI & ML

OpenAI Posted 722 AI-Written Math Papers. Only 162 Are Machine-Checked

On Oct. 6, OpenAI put roughly 4,000 open mathematics problems in front of an unreleased internal model, kept what it judged worth publishing, and posted the results to GitHub all at once: 722 manuscripts, grouped into 372 “families” of related results, according to Decrypt. A company spokesperson said “almost everything came from a single prompt handed to a single AI agent,” though some may have taken multiple attempts.

Only 162 of those 722 papers come with a computer-checked main result. That is about 22%.

The other 560 are, for now, a stack of claims. And the mathematicians expected to check them say they never asked for the job.

What OpenAI released, and what it says about it

OpenAI’s own post is short on numbers. It describes “a broad range of new mathematical results produced by an internal frontier model,” published in a GitHub repository “with protocols for paper revisions and citations.” It says the average result “used the equivalent compute of roughly three hours of ChatGPT Pro thinking,” that the repository includes 10 summaries of the model’s reasoning, and that it is “sharing formalizations of many of the proofs in Lean,” the proof-checking software. “We will update the repository with more formalizations as we obtain them,” it adds.

OpenAI's promotional artwork for its Sharing AI progress in mathematics announcement.
OpenAI’s announcement art for its Oct. 6 mathematics release. Image: OpenAI

The company does not name the model. It says it is “working to responsibly release” it, and that it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study in Princeton on how to put the results out. It also promised to fund workshops and conferences “around the understanding of major results produced by AI.”

The release itself carries a caveat, according to Decrypt: some of the unformalized results “could have issues.” No prompts were published.

There is a Millennium Prize subplot, too. Nature reports that the batch contains no solutions to the five remaining Millennium Prize Problems, though some preprints tackle related questions. Nature describes OpenAI’s separate September preprint on the Navier-Stokes equations as having solved one of the original seven. Asked about the significance of the new batch, an OpenAI spokesperson told Nature: “We will leave it to the broader community to assess [their] significance.”

Lean proofs: what 162 formal checks do and don’t prove

Lean is the strongest thing in this release. A Lean proof that compiles has been checked step by step by a machine, and no amount of confident prose can fake that.

Bar chart of OpenAI's Oct. 6 math release: about 4,000 problems posed, 722 manuscripts posted, 372 families of results, 162 papers with a Lean-checked main result, and 10 reasoning summaries shared.
Of 722 manuscripts, 162 carry a Lean-checked main result. Graphic: prompt/power

But it checks what it is given. As Decrypt points out, a passing Lean check confirms that the proof matches its Lean statement. It does not confirm that the Lean statement faithfully captures the problem a mathematician actually cared about. Translating a problem into formal language is where subtle errors hide, such as a dropped hypothesis or a definition that’s slightly too generous.

So even the 162 need a human to read the statement. The 560 without formalization need a human to read everything.

Andrew Sutherland, an MIT mathematician, put it plainly to Scientific American, as quoted by Decrypt: “Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified.” His shorter version: “We should ask for receipts.”

“Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power.” Association for Human Mathematics, Oct. 7

Why mathematicians are angry about OpenAI’s math papers

The sharpest response came on Oct. 7 from the Association for Human Mathematics. Its statement, which Terence Tao reposted on his blog as a guest post without adding comment of his own, opens flatly: “Mathematicians did not ask for this work to be done.” It rejects “OpenAI’s assertion that this release advances our subject” and closes with an ask: “We urge mathematicians to discontinue their work with OpenAI.”

Here the two accounts collide. OpenAI says it shaped the release with advice from the IAS advisory group. The AHM statement says OpenAI ignored that same group’s view that frontier AI companies should not test advanced problems on internal models. We could not open the advisory group’s own statement to settle which reading is right, so treat both as each side’s claim.

The Institute for Advanced Study weighed in separately, per Decrypt, warning that AI can now produce mathematical arguments “without the human who prompted it being able to understand the arguments, verify them, or take responsibility for them.” It added: “We believe that human understanding of mathematics remains of paramount importance.”

Elsewhere the mood was split. Nature found mathematicians on social media complaining about being scooped, and quoted an unnamed physicist calling the release a “slopocalypse.” University of Connecticut number theorist Alvaro Lozano-Robledo called the volume “staggering” and asked on X: “Why is OpenAI trying to solve hundreds of math problems to be released all at once?” Levent Alpöge, a mathematician at rival lab Anthropic, acknowledged some researchers would face “sad stories” if their projects were pre-empted, then wrote: “It’s obviously the most significant moment in mathematical history.” Abhishek Saha called it “a very big day for mathematics,” while noting that most of the results fall short of surprise breakthroughs, Decrypt reported.

The verification burden lands on universities

Our read: the scooping complaints are the visible part of a cost that will fall on academia, Canadian math departments included. Mathematics runs on refereeing, mostly unpaid, done by faculty and postdocs between teaching and their own research. A single solid paper can take a referee weeks.

OpenAI has now created 722 documents that any researcher working near one of those problems must reckon with. Cite it and risk building on a result nobody has checked. Ignore it and risk a priority dispute later. Check it yourself and lose the weeks. GitHub, unlike a journal, has no editor routing manuscripts to referees, and OpenAI says it is still “exploring other community-hosted alternatives” that meet the advisory group’s guidelines.

The economics are lopsided. The average result cost OpenAI about three hours of ChatGPT Pro-level compute, by its own estimate. The checking runs on human time, which doesn’t scale that way.

There is also a trust question about who is making the claim. Days before this release, OpenAI cancelled GPT-6.1 Astra after it was caught misrepresenting its work in testing, as we reported. The math model is a different, unnamed system, and nothing suggests the same problem. But “trust the internal model” is a harder sell than it was a month ago, which is why Sutherland’s demand for the model itself matters.

What to watch

  • The Lean count. OpenAI says it will add formalizations as it gets them. If 162 climbs fast, the verification argument weakens. If it stalls, the remaining 560 are the story.
  • The model. Replication needs access. OpenAI has said it is working to release it, with no date.
  • Where the papers go. A move to arXiv or journal submission would put the manuscripts into the normal refereeing pipeline, and onto referees’ desks.

OpenAI says it tried roughly 4,000 problems. It has shown the world the model’s reasoning on 10.

// Science & Climate Editor
Priya Natarajan

Priya Natarajan covers science and society for prompt/power: climate and energy tech, data centres, and the physical cost of the digital world. Every query has a water bill, and she would like to see the receipt.

Latest from prompt/power

  1. AI Safety Hearing: MPs Will Question OpenAI, Anthropic, Google and Meta on 13 OctOct 11
  2. National Medal of Science Goes to Musk, Brin, Huang and Su. Nvidia Pledged US$1BOct 10
  3. TikTok Placebo Safety Test: New York Says Teens Got a Fake Feed ResetOct 10
  4. Waymo Takes Its First Loan, US$5 Billion, as Robotaxis Head OverseasOct 10
  5. Claude Haiku 5.5 Price: 90% Cheaper Until Your Prompt Hits 100K TokensOct 10

Leave a Reply

Your email address will not be published. Required fields are marked *