
OpenAI's 722 Math Manuscripts: The Model Behind Them Has No Name, No Price and No API
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 127 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 56 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 320 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 54 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 347 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
At 21:47 UTC on October 6, 2026, a repository called openai/math appeared on GitHub with no description and an initial commit. Fourteen minutes later OpenAI posted it to X and put up a companion page, "Sharing AI progress in mathematics." Together those two things describe an event with no precedent: 722 mathematical manuscripts, grouped into 372 families, produced by an internal frontier model that OpenAI has not named, has not priced, and cannot be called from any API. The models it sits above are the ones you can actually reach today — GPT-6 Astra at $10 / $50 per million tokens and GPT-6.1 Sol at $2 / $10 — and neither of them wrote these papers. OpenAI says the model that did is still being trained, and that it is working on "responsibly releasing" it. This is what is verified so far, what is not, and the one question a reader should take from it.
Start with what this release is not. It is not a model launch: there is no model card, no identifier, no context window, no benchmark table, no date. It is not a paper: nothing here has been peer reviewed, and OpenAI says plainly that results without a Lean formalization "could have issues." And it is not a leak — the reason this piece is written in what-we-know-so-far framing is not that the material is secret, but that the subject itself, the model, is unreleased. Everything else below is checkable against the repository, OpenAI's own post, and the advisory group OpenAI says it consulted.
What actually landed on October 6
The repository is public, Apache-2.0 licensed, and written in Lean. Its README states the catalogue "contains 722 manuscripts organized into 372 families," where a family groups a principal result with companion arguments, consequences or alternative proofs. The manuscript map is detailed enough to audit: 372 families with a principal result each, listed with their constituent papers and, where available, a Lean formalization link.
OpenAI's post adds the provenance figures, which are its own and not independently audited:
• Problems posed — approximately 4,000, from which the released families were selected
• Compute per result — on average, the equivalent of roughly three hours of ChatGPT Pro thinking
• Reasoning summaries — 10 released, covering families including the irrationality exponent of π, the symmetric and general Mahler conjectures, Kaplansky's direct-finiteness conjecture in characteristic two, the Mézard–Parisi formula for diluted spin glasses and the three-dimensional relativistic Vlasov–Maxwell system
• Revision policy — new versions are posted for corrections, with earlier versions kept accessible
Subjects span the whole map: number theory, algebraic and complex geometry, combinatorics, theoretical computer science, operator algebras, probability and mathematical physics. Some of the family titles are on their face very large claims — a counterexample to Ryser's covering conjecture, integer multiplication below n log n, matrix multiplication with exponent at most 9/4, the Mahler conjectures, counterexamples to Baum–Connes and Kadison–Kaplansky. Whether a given one holds is exactly the question this release does not settle for you.

The model is unnamed, unreleased, and OpenAI says it is coming
The README is unambiguous about the author: "mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model," with the vast majority obtained "using an unreleased internal OpenAI model." No name, no codename, no size. Names circulating in other coverage have not been confirmed by OpenAI, and we are not going to supply one.
What is datable is the timeline around it. On August 28, 2026, per OpenAI's page on its mathematics advisory group, the company "began training a new internal model" that has since resolved the Navier–Stokes Millennium Prize problem and, by OpenAI's own count, more than 100 long-standing open problems. The Navier–Stokes result was announced separately on September 8, 2026, where OpenAI described the model behind it as "significantly more capable than GPT-6 Astra" and said training was still under way. Two of the manuscripts in this week's collection break the standard procedure, which is worth noting because it means the release is not one uniform pipeline: work on a zero-free region for the Riemann zeta function and a proof of the Hodge Conjecture for CM abelian varieties were handled differently, and OpenAI says the zeta write-up was human-edited for readability.
Then the sentence that matters for anyone planning around this: OpenAI writes that it is "working to responsibly release the model that produced these results." That is a statement of intent, not a date. Until it ships, the only OpenAI models a developer can call are the ones already on the price list.
What "verification" means here — and where it stops
This is the part most coverage of the release compresses into a phrase, and it is the part that decides how much weight to give any single paper.
OpenAI is explicit that "this collection includes results at different stages of verification." The strongest form of evidence in the bundle is a Lean formalization, which lets a computer check the proof rather than a person. The repository's own formalization manifest — the catalogue of papers with a formalized main result — lists 162 entries, so roughly one in five of the 722 manuscripts is backed by a machine-checkable proof in the public tree. OpenAI says more will be added "as we obtain them."
The remaining large majority is not verified in that sense, and OpenAI does not pretend otherwise over the whole collection: "Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly." Two structural details reinforce that this is a managed release rather than an open one — issues are disabled on the repository, and pull requests are restricted to collaborators. There is no public path to challenge a proof in place.
So the honest reading is narrower than the headline count suggests:
• Machine-checked — 162 manuscripts with a formalized main result in the public Lean tree
• Not machine-checked — the rest, which OpenAI itself flags as potentially containing errors
• Human responsibility — none named; OpenAI's own framing is that no person understands the arguments behind the bulk of what was produced
• Independent assessment — none published; the selection of which results were "significant" enough to include was made by OpenAI
None of that means the work is wrong. It means that for now, "OpenAI released 722 proofs" and "722 proofs have been checked" are different statements, and only the first one is true today.

The advisory group attached a condition, and it is on the record
OpenAI's post says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study and drew on "their advice and public recommendations." Those recommendations were published on September 29, 2026, after the group received more than 600 responses from the mathematical community, and they are worth reading directly because the group is not a rubber stamp.
The same document opens by stating that the group does not endorse the practice at all: "some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community... we do not endorse this practice, and we ask them to stop." The group also describes itself as operating independently, with members unpaid and no decision-making power over OpenAI — a description of the relationship that OpenAI's own post echoes.
Where the group's recommendations and this release line up is on process: a scoured literature review with citations to work that introduced the ideas, proof write-ups produced in a standard style rather than raw model output, and versioning that preserves earlier releases. Where they do not line up is on the part the group calls central — that a lab releasing results nobody understands should fund the human understanding that has to follow, through workshops, conferences and programs, and that this work remain community-led rather than lab-directed. OpenAI has committed to funding "a series of workshops, conferences, and special programs around the understanding of major results produced by AI" and says details will follow. That is the open item, and it is the one to hold them to.
What you can call today, and what routing is actually for here
There is no way to send a request to the model that produced these manuscripts, and no third-party platform can route it — any listing that claims otherwise is describing a model that does not exist in anyone's catalogue. The OpenAI models in production are the live GPT-6 line, and their numbers are on the price list rather than inferred:
• GPT-6 Astra — $10 / $50 per million tokens, $1 cache read; long-context tier above 272K prompt tokens at $20 / $75; 1.05M-token context, 128K max output
• GPT-6.1 Sol — $2 / $10 per million tokens, $0.10 cache read; same long-context tiering at $4 / $15; 1.05M-token context, 128K max output
Both are in the OrcaRouter catalogue, which is the practical point for anyone reading this release as a signal about capability: the models you are allowed to use today are also the models your prompts are already shaped around, and the gap between them and the internal one is precisely the gap OpenAI is asking mathematicians to help it close.
That gap is also a reason to keep the experimentation bill separate from the production path. When OpenAI does put a name and a price on this model, the first week will be a flood of unverified benchmark claims and a lot of panicked re-integration work — the routing argument for a platform like ours is that trying an unproven model then costs you a model string, not a migration. GPT-6 Astra sits at $10 / $50 and GPT-6.1 Sol is $2 / $10 per million tokens, passed through at list price with 0% markup, so a vendor price change is live on the same day it lands. Automatic failover means a new model can carry a fraction of traffic and get dropped without taking the request path down with it, and the routing DSL lets you compose several models into one call if you want a second opinion on a long derivation rather than a single answer. None of that applies to a model that has not shipped — but it is exactly the infrastructure you want in place before the next one does.

What to watch next
Four things, in the order they are likely to matter. The formalization count, because it is the only number in this release that can move from "OpenAI says" to "a machine checked it," and it is public. The first genuinely independent evaluation of a named result — a mathematician outside OpenAI working through a specific paper in public, not a summary of the collection. The community-hosted repository OpenAI says it is still exploring, which would move the artifact out of OpenAI's own account and into the field's hands. And the release of the model itself, which OpenAI has committed to but not scheduled.
Do the 722 manuscripts mean AI has solved these problems?
For a subset, if the mathematics holds: the collection claims counterexamples and proofs across well-known open problems, and 162 of the papers have a machine-checkable formalization behind their main result. For the rest, the claim is that a model produced a manuscript on the problem, which is a weaker statement than a verified solution, and OpenAI itself warns that unformalized results may contain errors. The distinction between "produced a proof" and "has a proof" is the whole story of this release.
Is any of this usable in a product today?
No. The manuscripts are PDFs, Lean sources and reasoning summaries — research artifacts, not a service. There is no model ID, no endpoint, no price and no access request form. The practically useful part of the release for a developer is the confirmation that OpenAI has a model materially beyond GPT-6 Astra in training, which is a planning signal for the next twelve months rather than something you can call this week.
Why did OpenAI release results instead of the model?
Because the two have different risk profiles. Publishing manuscripts is reversible — OpenAI keeps earlier versions and says it will fix issues — while shipping the model is not. OpenAI states it is working on releasing it responsibly, and the advisory group's recommendations, which the company says it drew on, are built around exactly this sequence: get the results into the field's hands with citations and versioning, then fund the human understanding that has to follow.
The one-line version for anyone skimming: a genuinely unprecedented body of AI-produced mathematics went public on October 6, roughly a fifth of it machine-checked, none of it peer reviewed, all of it attached to a model you cannot use. The interesting question is not whether the model is capable — three hours of frontier thinking compute per result across 4,000 problems answers that, and OpenAI says the model is significantly more capable than GPT-6 Astra — but whether the verification and the human understanding catch up before the next release makes this one obsolete.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
