The Daily Inference
AI & Technology · News · Developing

OpenAI's 722 maths papers are open. The model behind them is not.

The collection covers 372 families of results, including work on pi and major geometry conjectures. OpenAI has opened the papers for scrutiny while leaving the system behind them inaccessible, against its mathematical advisers' central recommendation.

Developing: this story is still unfolding and details may change.

OpenAI released 722 mathematical manuscripts on October 6, 2026, produced by an unnamed frontier model that the public cannot use. The collection includes work on how closely fractions can approximate pi, the Mahler conjectures in geometry and the Hodge Conjecture for a particular class of geometric objects. It puts hundreds of claimed results into circulation while leaving mathematicians without access to the system that generated them. [1] [2] [4]

The manuscripts are available in OpenAI's openai/math GitHub repository under Apache-2.0, a permissive open licence allowing anyone to reuse the material subject to its terms. They span pure mathematics, theoretical computer science and mathematical physics, grouped into 372 families of related results. Many, but not all, have accompanying proofs formalised in Lean, a programming language that allows computers to check mathematical arguments. [2] [4]

This is a developing release. The available sources establish no independent assessment of the full collection. OpenAI's own README warns: "Some of the unformalized results could have issues." [4]

The immediate dispute is about more than whether those arguments hold. On September 29, the Advisory Group on Mathematics and Artificial Intelligence, or AGMAI, asked labs to stop testing advanced mathematical problems on proprietary models. OpenAI says that advice informed the release, but has kept the model unavailable. The company says it is working to release it responsibly and has given no date. [1] [2] [6] [7]

Publishing the papers is useful. It lets mathematicians inspect the arguments instead of taking a company's word for them. But opening the output while withholding the tool divides the work unevenly: OpenAI controls access to the system, while the field must establish which results deserve to become accepted mathematics.

What is in the collection

The distinction between manuscripts and families matters. A family can contain a principal result, companion arguments, consequences or alternative proofs. The 722 documents therefore do not represent 722 distinct open problems solved. The repository organises the collection by mathematical discipline and gives each manuscript directory a BibTeX block, the standard citation information researchers can use when referring to it. [2] [3] [4]

OpenAI released 10 abridged summaries of the model's reasoning for selected results. These include work on the irrationality exponent of pi, which concerns how closely rational numbers can approximate it, and the symmetric and general Mahler conjectures, which concern the relationship between the volumes of a convex shape and its mathematical dual. [2] [4] These are named research problems, not merely exercises designed to test whether a model can reproduce a textbook answer.

Two results departed from the standard automated procedure. Work on a zero-free region for the Riemann zeta function was human-edited for readability. Such a region identifies where that function cannot equal zero, part of the mathematics used to study the distribution of prime numbers. The other exception was a manuscript presenting a proof of the Hodge Conjecture for CM abelian varieties, a particular class of geometric objects. The conjecture connects certain features of a space's topology with shapes defined by polynomial equations. [2] [4]

The qualification "for CM abelian varieties" is important. It describes the scope of the reported result, not a claim to have settled the Hodge Conjecture in full. Likewise, the selected reasoning summaries give readers a view of particular arguments, not a complete record of how all 372 families were produced. [4]

According to OpenAI, the model was posed approximately 4,000 open problems during evaluation. The company said the average result used computing resources equivalent to roughly three hours of ChatGPT Pro thinking, a comparison with processing time in its own product. That is not an estimate of how long a mathematician would need to understand the result. [2] [4]

Those figures alone do not establish a success rate. AGMAI recommended that labs releasing many results explain how the problems were chosen and how many others of comparable difficulty the model tried and failed to solve. [7] That record matters as much as the successes when assessing the system's capabilities. A collection of good answers can show achievement. It cannot, by itself, show reliability.

The repository also preserves its public release history. Corrections and revisions will appear as new versions while earlier versions remain accessible. [2] [4] That makes errors traceable rather than quietly erasable, a modest but essential feature of a scientific record.

Checking a proof is not testing a model

Lean is the release's strongest practical route to verification. Formalisation expresses a mathematical argument through explicit definitions and logical rules that software can check. Many of the manuscripts have accompanying Lean formalisations, though the available sources do not give a complete count. [2] [3] [4]

A proof does not become invalid because a machine wrote it. Nor does it become valid because a powerful company published it. The argument is what matters.

Computer checking strengthens that test, but researchers still need to establish that the formal statement matches the advertised result and uses the right assumptions. They also need to determine whether the result is new, what earlier work it depends on and why it matters. Logical correctness, novelty and credit are separate questions.

For manuscripts without formalisation, readers must inspect the written argument for gaps. That is familiar mathematical work. The unusual part is the volume arriving at once and the limited access to the process behind it.

The hidden model creates a second, distinct verification problem. Researchers can check a published proof without rerunning its author. But testing the model's broader capabilities requires trying other problems, varying prompts and examining failures. Published successes and selected reasoning summaries cannot substitute for that access.

OpenAI's stated intention to release the system responsibly addresses this criticism in principle. Without a date or public access, it does not yet address it in practice. [1] [2] The papers can earn acceptance result by result while claims about the system remain harder to test.

Nine weeks of rising stakes

The sequence began with OpenAI's release of 10 results on August 1. On September 8, the company announced a claimed solution to the Navier-Stokes Millennium Prize Problem. The October collection is its third major mathematics disclosure in roughly nine weeks. [5]

The September effort used approximately 10,000 concurrent AI agents and took 88 hours of compute, according to OpenAI. The company described the model as significantly more capable than GPT-6 Astra. [2] [5] Those figures describe a different scale of effort from the average attached to the new collection. They also underline why one headline number cannot describe the entire programme.

After the Navier-Stokes announcement, 25 Fields Medalists signed a declaration titled "A Severe Misalignment of AI in Mathematics". They objected to results arriving as press releases without named authors, clear attribution trails or time for the field to absorb them. [5] [8]

A dispute over credit made the concern concrete. NYU mathematician Tristan Buckmaster told TechCrunch that OpenAI pressured him not to credit collaborator Levent Alpöge, who works at Anthropic. OpenAI denied drawing directly on the pair's private work, while saying it could not rule out indirect benefit from de-identified usage data. This allegation and response concern the earlier work, not an established defect in the October manuscripts. [5]

Scientific American also reported that two results from the August release relied on existing published ideas without properly citing them, challenging OpenAI's assertion that the problems had seen no progress for at least a decade. [5] That history gives the field a reason to examine attribution alongside correctness.

The objections do not amount to rejection of machine-produced mathematics. Fields Medalist Timothy Gowers said he would recommend at least one earlier proof for a top journal without hesitation. [5] Good mathematics and a poor release process can coexist.

University of Toronto mathematician Daniel Litt put the case for publication plainly: "If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us." [1]

He is right. Withholding manuscripts would leave less to inspect. Publication is the beginning of scrutiny, however, not a substitute for it.

Advice adopted, access withheld

AGMAI was formed in September and is hosted at the Institute for Advanced Study. Its initial nine members were unpaid, could offer unsolicited advice and speak publicly, and controlled their own membership. They were expressly not responsible for advising OpenAI on the pace of its internal mathematical progress. [8]

After receiving more than 600 replies from the mathematical community, the group published its responsible-release recommendations on September 29. [2] [7] Its central position was blunt: "we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models." [7]

It also recommended scholarly repositories outside AI-lab control, disclosure of model names, prompts, reasoning records, time and computing costs, and formalisation where possible. It urged labs not to use mathematical releases as marketing vehicles and to fund human understanding through independent non-profit institutions. [7]

OpenAI has adopted useful parts of that programme: public manuscripts, many formalisations, a computing-cost estimate and a durable revision record. Its company-controlled repository, unnamed model and selective reasoning summaries stop short of the wider disclosure AGMAI sought. [2] [4] Following technical recommendations does not resolve the fundamental disagreement over access.

OpenAI spokesperson Lindsay McCallum Rémy said the advice informed the release and that the company would "continue incorporating community feedback and improving our standards for sharing major scientific advances". OpenAI also said it would fund workshops, conferences and special programmes to help people understand major AI-produced results. The available sources specify neither funding amounts nor a timetable. [1] [2] [6]

That support should now become concrete. Checking statements, tracing sources and explaining significant results are not incidental costs of producing mathematics. They are part of making the mathematics usable.

The next consequential developments will be independent assessments of particular families, documented corrections and funded work to understand the strongest results. A fuller record of failed attempts would make the capability claims testable. A model-release date would make OpenAI's promise measurable. The company has put 722 manuscripts on the table. Its next task is to give the people checking them more than papers.

Topics: OpenAI · AI agents · Peer review

Every edition in brief, three times a day, on our Telegram channel, on Bluesky and on Threads.

Sources
  1. OpenAI drops another batch of mathematical breakthroughs | The Verge The Verge
  2. OpenAI Releases 722 Math Manuscripts From an Unreleased AI Model - Unite.AI unite.ai
  3. OpenAI Publishes 722 AI-Generated Math Manuscripts, With Verification Still Uneven | Superpower Daily superpowerdaily.com
  4. GitHub - openai/math · GitHub github.com
  5. OpenAI drops 722 AI math proofs and mathematicians are not impressed - Startup Fortune startupfortune.com
  6. OpenAI Dumps 377 New Math Results on GitHub, Publishes Hand-Wringing Blog Post gizmodo.com
  7. general-sep29 - agmai.org agmai.org
  8. OpenAI forms math advisory group as its AI resolves more than 100 open problems | TechCrunch techcrunch.com