A Closer Look At OpenAI’s 722 AI Mathematics Proofs And What Comes Next
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: A Closer Look At OpenAI’s 722 AI Mathematics Proofs And What Comes Next on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI has published 722 mathematical manuscripts produced by an unreleased, unnamed model, covering 372 families of results drawn from roughly 4,000 problems. The manuscripts include claims about major open problems, but outside mathematicians have not yet confirmed them; some results also lack formal verification. The central test will be whether researchers can check, understand and build on the work.

OpenAI published 722 mathematical manuscripts on Monday, presenting results generated by an unnamed, unreleased model across 372 families of related results. The collection includes claims concerning prominent open problems, but OpenAI chief executive Sam Altman said the claims have not been confirmed by outside mathematicians, leaving verification and the mathematical value of the work unresolved.

According to OpenAI’s post and its GitHub repository, the manuscripts span number theory, geometry, topology, operator algebras, theoretical computer science and mathematical physics. The results were organized into 372 families after the model was given roughly 4,000 problems. OpenAI says it selected work it judged to have an appropriate level of significance; the source material says no outside group made that selection. The average result used about three hours of ChatGPT Pro reasoning compute, according to the source account.

The collection makes claims involving the Unique Games Conjecture, Hilbert’s tenth problem over the rationals, the isomorphism of nonabelian free group factors, a zero-free region for the Riemann zeta function to the right of Re(s) = 11/12, the Hodge conjecture for CM abelian varieties and the Mahler conjectures. These are claims made in the manuscripts, not results established by the release itself. OpenAI’s repository includes Lean formalizations for many, but not all, results. Its README cautions that some results without formalizations could have issues.

OpenAI also released 10 abridged reasoning summaries, drawn from the 372 families. The source account says two manuscripts followed exceptions to the usual process: the Riemann write-up was edited by humans for readability, and the Hodge result also received special treatment. The available information does not establish how much human intervention shaped the underlying arguments in each manuscript.

At a glance
reportWhen: Published Monday; independent review is…
The developmentOpenAI published 722 AI-generated mathematical manuscripts, including claims about several major open problems that have not been independently confirmed.
722 Proofs, One Question — Reality Check
AI Dispatch · Reality Check · 7 October 2026

722 proofs, one question: will any of OpenAI’s AI mathematics actually lead anywhere?

An unreleased, unnamed model produced claimed proofs of results that would each define a career. Sam Altman calls them “claims not yet confirmed by outside mathematicians.” The real question isn’t whether it’s impressive. It’s whether answers nobody understands become discoveries anyone can build on.

What was released
~4,000
problems posed to the model
→
372
families judged significant — by OpenAI
→
722
manuscripts, Apache-2.0, GitHub
·
10
reasoning summaries — for 372 families
Average result: ~3 hours of ChatGPT Pro thinking compute. Lean formalizations for many, not all. OpenAI’s README: “some of the unformalized results could have issues.”
A sample of what’s claimed — any one would define a career
Unique Games Conjecture
The central open problem in hardness of approximation.
LEAN · reported
Quasi-Riemann hypothesis
Zeta has no zeros with Re(s) > 11/12. Exception to the standard procedure; write-up human-edited.
LEAN · reported
Free group factors are isomorphic
Open since the 1940s; central to operator algebras.
LEAN · reported
Hilbert’s tenth problem over ℚ
Is there an algorithm deciding rational solutions?
STATUS · see repo
Hodge for CM abelian varieties
A special case of the Hodge conjecture, itself a Millennium Prize problem. Exception to the standard procedure.
STATUS · see repo
Mahler conjectures
Symmetric and general cases, convex geometry.
STATUS · see repo
None independently confirmed. Lean-checked doesn’t mean the formal statement matches the conjecture mathematicians mean — see below.
The track record so far — the first three releases tell you most of what to expect from the fourth
May 2026
Erdős unit distance
HELD UP

Same day: Alon, Bloom, Gowers, Litt, Sawin post a digested, human-verified version. The model for success.

Aug 2026
“Ten Advances”
ONE DISPUTED

Connes rigidity counterexample challenged within a day — constructed groups fail the required condition. Three rival machine “counterexamples” from different labs now circulate.

Sep 2026
Navier–Stokes
LEAN-CHECKED · CONTESTED

~10,000 agents, 88 hours, est. ~$22M at retail. Priority dispute; 25 Fields Medalists sign “A Severe Misalignment” — not saying it’s wrong, saying it’s not understood.

Oct 2026
722 manuscripts
UNVERIFIED

Altman now hedges at announcement — a shift from September. Verification has barely started.

Three fates for every AI proof — and only one of them is a discovery
① Digested
A new idea others use

Humans extract the technique, write it up, build on it. This is where downstream discovery comes from.

Like: Wiles → modularity · Perelman → Ricci flow surgery · Erdős counterexample, May 2026
② Settled but sterile
True, checked, unexplained

The question is answered; nobody learns anything reusable. Closes a door without opening a field.

Like: the Four Colour Theorem (1976) — a computer case-check that produced comparatively little new theory
③ Wrong, or wrong thing
Fails, or proves a near-miss

The proof breaks, or proves a statement that doesn’t match the conjecture as mathematicians mean it.

Like: the disputed Connes counterexample, August 2026
Which bucket each of the 372 families lands in isn’t a question about the AI. It’s a question about whether humans do the work of understanding it.
✓ Where downstream value is real — a literature is waiting
A literature of results “assuming UGC”— if proved →Theorems overnight

The Unique Games Conjecture is the clearest case. Results like the optimality of Goemans–Williamson for Max-Cut are proved assuming UGC. A correct proof converts them all — no understanding required. A zero-free strip for zeta works the same way for prime-distribution results. Free group factors, Kadison, Mahler would redirect whole programmes — but how depends on the method, which means digestion.

✕ What not to expect

Technology. A Navier–Stokes blow-up proof doesn’t change how anyone designs aircraft; engineering turbulence models never depended on the answer. Near-term consequences are mathematical, not industrial. “AI will cure cancer next” skips several steps.

◆ The real bottleneck: adjudication, not proof
Lean checksThe proof follows from the formal statement
but
Lean doesn’t checkWhether the formal statement is the conjecture
so
Still needsA human expert, per result — and the field has a fixed supply of them

“Verification abundance, adjudication scarcity” — making proof-checking cheap doesn’t reduce the burden of deciding what’s true and what matters. 722 manuscripts land on a review system built for a trickle, filtered by a selection nobody outside OpenAI made.

What the IAS advisory group asked for — and what OpenAI did
The group asked for
OpenAI’s release
Status
Repository not controlled by an AI lab
OpenAI’s GitHub; “exploring” alternatives
NO
Name of the model
Unnamed internal model
NO
Prompts used
Not published
NO
Summarized chain of thought per result
10 summaries for 372 families
PARTIAL
Time and compute cost
~3 hours Pro compute on average
YES
How many problems tried and failed
~4,000 posed; per-problem detail not in README
PARTIAL
Formalization where possible
Many, not all
PARTIAL
Funding for understanding, via existing non-profits
Workshops promised; mechanism unspecified
PARTIAL
The group’s recommendations open with a line OpenAI’s post doesn’t quote: it does not endorse labs testing advanced problems on proprietary models, and asks them to stop. Real progress over September — still short on the items that matter most for adjudication.
Signals that will tell you whether discovery is happening
01
Digest papers

Humans re-deriving results, like Alon–Gowers et al. in May

02
Citations

Other people’s work building on these manuscripts

03
Errata rate

How many unformalized results survive expert checking

04
Statement audits

Do the Lean statements match the real conjectures?

05
Journals

Do any survive peer review?

The take

Some of it, yes — where a literature is waiting (UGC), a correct proof pays off immediately; where a proof carries a new technique humans digest, it can open a field. Most of it, probably not on its own: at 722 manuscripts with 10 reasoning summaries, the Four Colour pattern is the likely default unless mathematicians are funded and given time. And some will be wrong — OpenAI says so itself. It’s an industry pattern, not one company’s: the forced-Euler result came from an Anthropic researcher, and rival machine-generated Connes “counterexamples” circulate from different labs. The proofs arrived this week. The discoveries, if they come, will arrive at the speed of human understanding.

Sources: OpenAI, “Sharing AI progress in mathematics” (6 Oct 2026) and openai/math README; catalogue contents via OfficeChai & AI Daily Digest; OpenAI Navier–Stokes post (8 Sep 2026); ~$22M estimate attributed to Zvi Mowshowitz via arXiv:2609.28591; Erdős and Connes history via arXiv:2608.28997; Fields Medalists’ declaration (11 Sep 2026); AGMAI “Responsible Release of AI-Generated Mathematics” (29 Sep 2026). No catalogue claim independently verified here. Lean status per reporting. Not investment advice.
thorstenmeyerai.com

Verification Will Shape the Impact

The number of manuscripts and the prominence of some claims make the release a substantial test of how AI-generated mathematics enters research. But a result’s importance is not settled by its headline or by the volume of work published. Mathematicians must determine whether each argument is correct, whether it proves the stated result, and whether its methods can be understood and reused.

The source account points to OpenAI’s May result on the Erdős unit-distance conjecture as one possible model: mathematicians produced a digested, human-verified version of the machine-generated work. That distinction matters because verification is not the same as understanding. A checked proof may settle a question, while a proof whose ideas researchers can extract may also lead to further results. Until independent researchers inspect this catalogue, its broader contribution remains uncertain.

The opposite outcome is also possible. A manuscript may contain an error, or establish a statement different from the conjecture researchers intended to test. Even a correct proof might offer little reusable insight. The practical consequence for readers and researchers is that publication is a starting point for review, not evidence that dozens of longstanding problems have been settled.

Amazon

mathematics proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

OpenAI’s Earlier Math Releases

This is described in the source material as OpenAI’s fourth major mathematics release of the year. In May, the company’s model produced a counterexample to the Erdős unit-distance conjecture; five mathematicians then posted a human-verified account. In August, OpenAI announced “Ten Advances.” One claimed counterexample to Connes’s rigidity conjecture was disputed within a day, with a critique arguing that the constructed groups did not meet the conjecture’s required condition.

In September, OpenAI announced a Lean-formalized result about finite-time blow-up in the Navier–Stokes equations, produced with about 10,000 concurrent agents over 88 hours, according to the source account. That announcement also prompted debate about the priority of related work and about using famous problems as AI benchmarks. A declaration signed by 25 Fields Medalists criticized the approach as misaligned with the aims of mathematics. Their objection, as described in the source, focused on human understanding and research practice, rather than asserting that the proof was wrong.

These episodes show why the present release needs to be judged result by result. Earlier work has included a result that mathematicians verified, as well as a claim that drew a prompt technical objection. Neither precedent confirms or disproves the new manuscripts. The status of each claim depends on independent scrutiny, not on the model’s record taken as a whole.

Amazon

AI mathematical research tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Which Proofs Will Survive Review?

No outside confirmation of the 722 manuscripts is reported in the source material. It does not identify which of the 372 families have been reviewed by independent specialists, which have been formally checked, or whether any review has uncovered errors. The repository’s warning specifically leaves open problems in some manuscripts without formalizations.

OpenAI has not named or released the model, and the available account does not provide a full description of how it was trained or how its reasoning was supervised. Nor does it explain the criteria in enough detail for outsiders to reproduce the company’s selection of significant work from roughly 4,000 prompts. The claims may vary widely in strength and importance; the collection-wide count cannot establish the validity of individual results.

It is also too early to know whether the proofs will yield methods that other mathematicians can apply. The source material identifies the Unique Games claim as potentially consequential because of its links to theoretical computer science, but it does not report a verified proof or explain how the relevant literature would change if the claim held. That impact remains conditional on both correctness and expert interpretation.

Amazon

formal proof assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Checking Comes Next

The next step is for specialists to examine the manuscripts, test their arguments and, where possible, check formalizations. That work will need to distinguish a correct proof of the stated claim from an argument that proves a nearby but different statement. The collection’s scale means the review is likely to proceed manuscript by manuscript, rather than through one verdict on all 722 papers.

Researchers will also need to translate any sound results into explanations that the field can assess and use. The earlier Erdős example, in which mathematicians prepared a digested version, offers one precedent described in the source material. Whether the new papers produce comparable accounts, lead to further theorems or are revised after scrutiny has not yet been reported.

For now, the release establishes that OpenAI is presenting a large body of AI-generated mathematical work—not that its major claims are settled. The clearest milestones to watch are independent evaluations, corrections or confirmations, along with evidence that researchers can build on the methods. Until those appear, the catalogue is a research program under review, not a verified set of discoveries.

Amazon

mathematical problem solving software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI publish?

OpenAI published 722 mathematical manuscripts, organized into 372 families of related results and drawn from roughly 4,000 problems posed to its model, according to the source material.

Has the mathematics been independently confirmed?

Not according to the information provided. Altman described the results as claims not yet confirmed by outside mathematicians. The manuscripts require independent review, and some are not formally verified.

What major problems do the manuscripts address?

The collection includes claims about the Unique Games Conjecture, Hilbert’s tenth problem over the rationals, the Hodge conjecture for CM abelian varieties and other problems. Their inclusion does not establish that the claims are correct.

Why does formal verification matter?

Formalization can let proof-checking software test whether a proof follows from its stated assumptions. OpenAI’s repository includes Lean formalizations for many, but not all, results, and warns that some unformalized results could have issues.

What happens after publication?

Mathematicians must inspect the arguments, verify what each manuscript actually proves and assess whether its methods can support further research. The source material reports no collection-wide independent verdict yet.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

OlmoEarth Studio’s Embedding Exports: Boosting AI Downstream Tasks

OlmoEarth Studio now supports on-demand generation and export of satellite data embeddings, aiding downstream AI tasks like similarity search and land classification.

Partial Lunar Eclipse Visible From Barcelona On August 28

A partial lunar eclipse will be visible from Barcelona on August 28, with details on timing and appearance confirmed by local astronomical sources.

Evidence Of Fraud In An Influential Study About Procrastination

Investigations reveal potential misconduct in a widely-cited research on procrastination, raising questions about its findings and impact.

Masters of Doubt: How Arcesilaus and Carneades Led the Skeptical Academy

Masters of doubt, Arcesilaus and Carneades reshaped philosophy through relentless questioning, leaving us to wonder how their skeptical legacy endures today.