OpenAI’s unreleased model wrote 722 math papers
On October 6, OpenAI published 722 math papers on GitHub, all produced by an internal model it hasn’t released yet. The papers are grouped into 372 result families spanning 17 fields, from number theory to partial differential equations.
None of them has been peer reviewed. OpenAI itself acknowledges that some of the results that haven’t been formally verified could have issues.
The claims are not small. The catalog includes claimed solutions to questions that have been open for decades, such as the quasi-Riemann hypothesis, a weaker form of the Riemann hypothesis, Hilbert’s tenth problem over the rational numbers, and whether the irrationality exponent of π is exactly 2.
OpenAI says the model was given about 4,000 problems. The catalog is made up of the results it judged significant enough.
OpenAI says the vast majority of results came from the same procedure, using on average the equivalent of three hours of ChatGPT Pro thinking per result. A spokesperson told Scientific American that almost all of them came from a single prompt given to a single agent, though some may have taken several attempts.
The prompt in one of the reasoning summaries OpenAI shared starts like this: “Even if the problem is ‘open’, the intention is that you should resolve it and present a full solution.”
The first error surfaced a day after the release. On October 7, OpenAI withdrew one paper because of a sign error, along with two papers that depended on it, and revised 14 others. The catalog now has 719 papers.
At release, only 162 papers had their main result verified in Lean, a language that lets computers check proofs. OpenAI now says about 42% of the top-line results are formalized in Lean.
According to OpenAI, the same model was behind an even bigger claim in September. On September 8, the company said about 10,000 agents running on this model solved the Navier-Stokes Millennium Prize Problem in 88 hours. A chart in the company’s announcement shows the model well ahead of GPT-6 Astra on a curated set of open problems.
That result didn’t come from a single prompt: the agents were first given a result found for the Euler equations, and their intermediate findings were merged with Codex.
The announcement also sparked a priority dispute. OpenAI says its effort began on September 1 after a rumor, which it later realized was about NYU’s Tristan Buckmaster and Anthropic employee Levent Alpöge. The pair had solved the forced version of the Euler equations.
Buckmaster wrote that he asked whether the model had been trained on their Codex sessions and got no answer, adding that he wasn’t accusing anyone. After an investigation, OpenAI said his prompts from the prior two months could not have influenced the system.
Verification is still ongoing. On September 11, the Clay Mathematics Institute said the problem had “apparently been settled” but stressed that its process is “deliberately unhurried.” For now, it lists Navier-Stokes under “Active problems,” separate from both solved and unsolved.
The 166-page proof hasn’t been published in a peer-reviewed journal yet, and OpenAI says it won’t claim the prize. One preprint argues the Lean proof doesn’t fully match the written one.
Much of the criticism is about transparency. In September, Terence Tao criticized “the refusal of AI companies to disclose their negative results, or reveal the process towards obtaining their solutions.” The IAS advisory group OpenAI consulted had recommended disclosing the prompts, time and cost for each result.
OpenAI shared average compute and 10 reasoning summaries, but no per-result prompts or times.
Sources
- OpenAI, “Sharing AI progress in mathematics”, (openai.com)
- OpenAI (GitHub), GitHub repository (openai/math, README), (github.com)
- OpenAI (GitHub), “History”, (github.com)
- OpenAI (GitHub), “OpenAI Research Catalog”, (github.com)
- OpenAI (GitHub), “The Mézard–Parisi formula for diluted spin glasses”, reasoning summary (PDF), (github.com)
- Scientific American, “OpenAI unleashes hundreds more math results upon a field already in shock”, (scientificamerican.com)
- OpenAI, “On the Navier–Stokes Millennium Prize Problem”, (openai.com)
- OpenAI, “Finite Time Blowup for Navier–Stokes”, proof writeup (PDF), (cdn.openai.com)
- OpenAI, “Advisory Group on Mathematics and Artificial Intelligence”, (openai.com)
- Tristan Buckmaster (NYU), statement (PDF), (cims.nyu.edu)
- Clay Mathematics Institute, “Navier-Stokes Announcement”, (claymath.org)
- Clay Mathematics Institute, “The Millennium Prize Problems” (claymath.org)
- Terence Tao (Mastodon), Mastodon post (part 3 of a 4-part thread), (mathstodon.xyz)
- Advisory Group on Mathematics and AI (AGMAI), “Responsible Release of AI-Generated Mathematics”, (agmai.org)
- arXiv (Bastounis, Circelli, Hansen), “Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs”, (arxiv.org)
About this story
This story was posted on Instagram by @jarrus.tech on Oct. 8, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/c27
This story in Turkish: OpenAI’ın modeli 722 matematik makalesi yazdı
On the same topic
DeepMind-born drug discovery company Isomorphic Labs in funding talks at a $40-50 billion valuation
Isomorphic Labs, the DeepMind-born AI drug discovery company, is in early talks to raise money at a valuation of at least $40 billion, Bloomberg reported.
Anthropic releases Claude Haiku 5.5 at 90% lower prices
Claude Haiku 5.5, released Oct. 7, costs 90% less than Haiku 4.5 for prompts up to 100,000 tokens. It leads on benchmarks but uses more tokens per task.
OpenAI is rolling out GPT-6 to everyone in ChatGPT
OpenAI began rolling out GPT-6 to everyone in ChatGPT on October 7. Its new Intelligent UI builds answers from text, visuals and interactive elements.