🔍 Read the full analysis: A Closer Look At OpenAI’s 722 AI Mathematics Proofs And What Comes Next on ThorstenMeyerAI.com
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
OpenAI has published 722 mathematical manuscripts produced by an unreleased, unnamed model, covering 372 families of results drawn from roughly 4,000 problems. The manuscripts include claims about major open problems, but outside mathematicians have not yet confirmed them; some results also lack formal verification. The central test will be whether researchers can check, understand and build on the work.
OpenAI published 722 mathematical manuscripts on Monday, presenting results generated by an unnamed, unreleased model across 372 families of related results. The collection includes claims concerning prominent open problems, but OpenAI chief executive Sam Altman said the claims have not been confirmed by outside mathematicians, leaving verification and the mathematical value of the work unresolved.
According to OpenAI’s post and its GitHub repository, the manuscripts span number theory, geometry, topology, operator algebras, theoretical computer science and mathematical physics. The results were organized into 372 families after the model was given roughly 4,000 problems. OpenAI says it selected work it judged to have an appropriate level of significance; the source material says no outside group made that selection. The average result used about three hours of ChatGPT Pro reasoning compute, according to the source account.
The collection makes claims involving the Unique Games Conjecture, Hilbert’s tenth problem over the rationals, the isomorphism of nonabelian free group factors, a zero-free region for the Riemann zeta function to the right of Re(s) = 11/12, the Hodge conjecture for CM abelian varieties and the Mahler conjectures. These are claims made in the manuscripts, not results established by the release itself. OpenAI’s repository includes Lean formalizations for many, but not all, results. Its README cautions that some results without formalizations could have issues.
OpenAI also released 10 abridged reasoning summaries, drawn from the 372 families. The source account says two manuscripts followed exceptions to the usual process: the Riemann write-up was edited by humans for readability, and the Hodge result also received special treatment. The available information does not establish how much human intervention shaped the underlying arguments in each manuscript.
722 proofs, one question: will any of OpenAI’s AI mathematics actually lead anywhere?
An unreleased, unnamed model produced claimed proofs of results that would each define a career. Sam Altman calls them “claims not yet confirmed by outside mathematicians.” The real question isn’t whether it’s impressive. It’s whether answers nobody understands become discoveries anyone can build on.
Same day: Alon, Bloom, Gowers, Litt, Sawin post a digested, human-verified version. The model for success.
Connes rigidity counterexample challenged within a day — constructed groups fail the required condition. Three rival machine “counterexamples” from different labs now circulate.
~10,000 agents, 88 hours, est. ~$22M at retail. Priority dispute; 25 Fields Medalists sign “A Severe Misalignment” — not saying it’s wrong, saying it’s not understood.
Altman now hedges at announcement — a shift from September. Verification has barely started.
Humans extract the technique, write it up, build on it. This is where downstream discovery comes from.
The question is answered; nobody learns anything reusable. Closes a door without opening a field.
The proof breaks, or proves a statement that doesn’t match the conjecture as mathematicians mean it.
The Unique Games Conjecture is the clearest case. Results like the optimality of Goemans–Williamson for Max-Cut are proved assuming UGC. A correct proof converts them all — no understanding required. A zero-free strip for zeta works the same way for prime-distribution results. Free group factors, Kadison, Mahler would redirect whole programmes — but how depends on the method, which means digestion.
Technology. A Navier–Stokes blow-up proof doesn’t change how anyone designs aircraft; engineering turbulence models never depended on the answer. Near-term consequences are mathematical, not industrial. “AI will cure cancer next” skips several steps.
“Verification abundance, adjudication scarcity” — making proof-checking cheap doesn’t reduce the burden of deciding what’s true and what matters. 722 manuscripts land on a review system built for a trickle, filtered by a selection nobody outside OpenAI made.
Humans re-deriving results, like Alon–Gowers et al. in May
Other people’s work building on these manuscripts
How many unformalized results survive expert checking
Do the Lean statements match the real conjectures?
Do any survive peer review?
Some of it, yes — where a literature is waiting (UGC), a correct proof pays off immediately; where a proof carries a new technique humans digest, it can open a field. Most of it, probably not on its own: at 722 manuscripts with 10 reasoning summaries, the Four Colour pattern is the likely default unless mathematicians are funded and given time. And some will be wrong — OpenAI says so itself. It’s an industry pattern, not one company’s: the forced-Euler result came from an Anthropic researcher, and rival machine-generated Connes “counterexamples” circulate from different labs. The proofs arrived this week. The discoveries, if they come, will arrive at the speed of human understanding.
Verification Will Shape the Impact
The number of manuscripts and the prominence of some claims make the release a substantial test of how AI-generated mathematics enters research. But a result’s importance is not settled by its headline or by the volume of work published. Mathematicians must determine whether each argument is correct, whether it proves the stated result, and whether its methods can be understood and reused.
The source account points to OpenAI’s May result on the Erdős unit-distance conjecture as one possible model: mathematicians produced a digested, human-verified version of the machine-generated work. That distinction matters because verification is not the same as understanding. A checked proof may settle a question, while a proof whose ideas researchers can extract may also lead to further results. Until independent researchers inspect this catalogue, its broader contribution remains uncertain.
The opposite outcome is also possible. A manuscript may contain an error, or establish a statement different from the conjecture researchers intended to test. Even a correct proof might offer little reusable insight. The practical consequence for readers and researchers is that publication is a starting point for review, not evidence that dozens of longstanding problems have been settled.
mathematics proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
OpenAI’s Earlier Math Releases
This is described in the source material as OpenAI’s fourth major mathematics release of the year. In May, the company’s model produced a counterexample to the Erdős unit-distance conjecture; five mathematicians then posted a human-verified account. In August, OpenAI announced “Ten Advances.” One claimed counterexample to Connes’s rigidity conjecture was disputed within a day, with a critique arguing that the constructed groups did not meet the conjecture’s required condition.
In September, OpenAI announced a Lean-formalized result about finite-time blow-up in the Navier–Stokes equations, produced with about 10,000 concurrent agents over 88 hours, according to the source account. That announcement also prompted debate about the priority of related work and about using famous problems as AI benchmarks. A declaration signed by 25 Fields Medalists criticized the approach as misaligned with the aims of mathematics. Their objection, as described in the source, focused on human understanding and research practice, rather than asserting that the proof was wrong.
These episodes show why the present release needs to be judged result by result. Earlier work has included a result that mathematicians verified, as well as a claim that drew a prompt technical objection. Neither precedent confirms or disproves the new manuscripts. The status of each claim depends on independent scrutiny, not on the model’s record taken as a whole.
As an affiliate, we earn on qualifying purchases.
Which Proofs Will Survive Review?
No outside confirmation of the 722 manuscripts is reported in the source material. It does not identify which of the 372 families have been reviewed by independent specialists, which have been formally checked, or whether any review has uncovered errors. The repository’s warning specifically leaves open problems in some manuscripts without formalizations.
OpenAI has not named or released the model, and the available account does not provide a full description of how it was trained or how its reasoning was supervised. Nor does it explain the criteria in enough detail for outsiders to reproduce the company’s selection of significant work from roughly 4,000 prompts. The claims may vary widely in strength and importance; the collection-wide count cannot establish the validity of individual results.
It is also too early to know whether the proofs will yield methods that other mathematicians can apply. The source material identifies the Unique Games claim as potentially consequential because of its links to theoretical computer science, but it does not report a verified proof or explain how the relevant literature would change if the claim held. That impact remains conditional on both correctness and expert interpretation.
As an affiliate, we earn on qualifying purchases.
Independent Checking Comes Next
The next step is for specialists to examine the manuscripts, test their arguments and, where possible, check formalizations. That work will need to distinguish a correct proof of the stated claim from an argument that proves a nearby but different statement. The collection’s scale means the review is likely to proceed manuscript by manuscript, rather than through one verdict on all 722 papers.
Researchers will also need to translate any sound results into explanations that the field can assess and use. The earlier Erdős example, in which mathematicians prepared a digested version, offers one precedent described in the source material. Whether the new papers produce comparable accounts, lead to further theorems or are revised after scrutiny has not yet been reported.
For now, the release establishes that OpenAI is presenting a large body of AI-generated mathematical work—not that its major claims are settled. The clearest milestones to watch are independent evaluations, corrections or confirmations, along with evidence that researchers can build on the methods. Until those appear, the catalogue is a research program under review, not a verified set of discoveries.
mathematical problem solving software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI publish?
OpenAI published 722 mathematical manuscripts, organized into 372 families of related results and drawn from roughly 4,000 problems posed to its model, according to the source material.
Has the mathematics been independently confirmed?
Not according to the information provided. Altman described the results as claims not yet confirmed by outside mathematicians. The manuscripts require independent review, and some are not formally verified.
What major problems do the manuscripts address?
The collection includes claims about the Unique Games Conjecture, Hilbert’s tenth problem over the rationals, the Hodge conjecture for CM abelian varieties and other problems. Their inclusion does not establish that the claims are correct.
Why does formal verification matter?
Formalization can let proof-checking software test whether a proof follows from its stated assumptions. OpenAI’s repository includes Lean formalizations for many, but not all, results, and warns that some unformalized results could have issues.
What happens after publication?
Mathematicians must inspect the arguments, verify what each manuscript actually proves and assess whether its methods can support further research. The source material reports no collection-wide independent verdict yet.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
