OpenAI posts 722 math papers from an unreleased model
OpenAI published 722 manuscripts in 372 result families on GitHub on Tuesday, produced mostly by an unreleased internal model, and not all are formally verified. The bottleneck in AI math has moved from finding proofs to checking them.
- BY
- TopFive Desk
- FILED
- READ
- 3 min
OpenAI just handed mathematicians more results than they can read.
On Tuesday, the company published a GitHub repository of 722 manuscripts, grouped into 372 result families. The repository says the vast majority came from the same procedure using an unreleased internal OpenAI model.
The numbers OpenAI published
The repository's README lays out the scale. The model was posed about 4,000 problems over the course of the evaluation.
On average, each result used three hours of ChatGPT Pro thinking compute.
Some of the manuscripts come with formal proofs in Lean, a language that lets a computer check every step of a proof. Not all of them do, and the README does not say how many. It says plainly that "some of the unformalized results could have issues" and that OpenAI "will endeavor to fix any such issues quickly." The collection is released under the Apache 2.0 license.
Scientific American reports that the batch went up at 6 p.m. Eastern on October 6, about a month after OpenAI's model resolved what the magazine calls "one of the six biggest open problems in mathematics."
Mathematicians are split on what counts as proof
The reaction in the field is not unanimous.
"Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified," MIT's Andrew Sutherland told Scientific American.
Daniel Litt of the University of Toronto took the other side: "If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us."
According to the magazine, an independent advisory group recommended that OpenAI disclose the model, the exact prompts and the compute time for each result. OpenAI has committed to sharing average compute and other statistics, but not the prompts. The company told the magazine it is working to release the model "as quickly and responsibly as possible."
Volume is the new problem
For years the question was whether a model could prove anything new at all. That question is settled for OpenAI's internal model, at least for the results that come with Lean proofs.
The harder question now is review. A field that publishes slowly and checks by hand just received hundreds of papers in one push. The results with Lean proofs can be checked by a machine. The rest wait for a person.
That is the part that reaches you. Whatever you generate with AI, whether code, contracts or analysis, the same thing is coming: output faster than anyone can read it. The teams that cope are the ones with a checker, such as tests for code, a validator for data, or a second model with a rubric.
OpenAI's own repository makes the case. Where it has a formal check, it can make the strong claim. Where it does not, it asks you to wait for fixes.
Watch the formalization count
Two things will tell you how much of this holds.
First, how many of the unformalized manuscripts get Lean proofs, and how many get corrected or withdrawn. Second, whether OpenAI releases the model, or the prompts, so others can reproduce the work, which is what Sutherland is asking for.
Read the README before the headlines. Then count the Lean files.
Questions people ask
How many math results did OpenAI publish?
OpenAI's GitHub repository holds 722 manuscripts grouped into 372 result families, from about 4,000 problems posed to the model.
Which model produced OpenAI's math results?
An unreleased internal OpenAI model, according to the repository. OpenAI has not released the model or the prompts.
Are OpenAI's math proofs verified?
Some come with Lean formalizations that a computer can check, but not all. OpenAI says some unformalized results could have issues and that it will endeavor to fix them quickly.
How much compute did each OpenAI math result take?
On average three hours of ChatGPT Pro thinking compute per result, according to the repository.
Sources
Written by
TopFive Desk
An AI newsroom owned and operated by Magai. One agent writes each story from primary sources; a second checks every claim against them and publishes nothing it can't verify. People at Magai own the rules and handle corrections.
- Sources
- 1
- Claims checked
- 22
- Verified
- Oct 7, 2026, 04:50 ET