OpenAI has published 377 results from a private advanced model across mathematics, including work on unsolved problems, while researchers demand clearer evidence of how the proofs were produced and independently verified.
OpenAI has released 377 mathematical results generated by an advanced model that the public cannot access. The collection reaches into algebra, number theory, theoretical computer science, mathematical logic and topology, and includes work on questions that remain unsolved at the frontier of research.
The publication, dated 6 October 2026, consists of 722 manuscripts grouped into 372 families of results in a public GitHub repository. OpenAI says the internal frontier model examined about 4,000 mathematical problems, with an average attempt requiring roughly three hours of computation. The repository also includes estimates of computational resources and statistics on successful and unsuccessful attempts.
The publication is not simply a capability announcement. It is also a test of whether a machine-produced proof can become usable mathematical knowledge when the system behind it is closed and the contribution of human researchers is difficult to measure.
OpenAI said the average result required about three hours of computing. Many of the proofs were checked with Lean, a computer language used to verify that the logical steps in a formal proof are valid. The release does not state that every result has a Lean formalization. That distinction matters: Lean can establish that a formalized argument follows its specified rules, but it does not by itself explain how the model found the argument, determine whether the underlying idea is novel or replace independent mathematical review.
The company has provided OpenAI's release details, including protocols for correcting manuscripts and updating references. Those procedures acknowledge that the published claims remain open to checking and revision rather than constituting completed academic verification of all 377 results.
For 10 solutions the company supplied summaries of the model's route to its answers and said that additional details about how the results were obtained would be published. Those summaries offer additional context, but they are not equivalent to opening the model for independent examination. The model itself remains unavailable to the public, and its parameters and conditions for independent reproduction have not been provided.
The released material also does not establish that every result represents a wholly new mathematical idea generated without assistance from existing research. It does not disclose a complete record of prompts, intermediate attempts, human interventions or training data relevant to each result.
The results emerged from internal testing rather than a public product. Dan Roberts, OpenAI's research lead, described the proofs as a byproduct of testing models intended to help develop better tools. That framing places the release closer to a company research disclosure than to an independently audited benchmark.
OpenAI's earlier mathematical evaluations focused on established challenges such as problems used in the International Mathematical Olympiad. After its models were able to solve those tasks, the company developed tests based on unsolved questions at the edge of mathematical research. Some of the 377 released results came from that newer category.
This shift changes what a successful answer means. A competition problem normally has a known solution and a defined judging process. An unsolved research question requires specialists to establish that the statement is correct, that the proof is complete and that the result genuinely advances the field. A system can produce a correct formal derivation while leaving the intellectual origin and broader significance of the result unresolved.
That distinction is familiar across research disciplines. Whether a result appears in a mathematics repository, a computational archive or a journal such as Nature, specialists must still inspect the assumptions, compare the claim with prior literature and assess whether the argument changes the field. Formal checking strengthens one part of that process, but it does not answer every question about discovery or significance.
The release follows OpenAI's announcement last month that it had solved the Navier-Stokes problem. OpenAI's claim concerned one of the seven Millennium Prize Problems identified by the Clay Mathematics Institute. Each carries a $1 million reward for a solution; the institute describes the formal requirements and review pathway for these problems in its Millennium Prize framework. A company announcement is not the same as the independent mathematical process required to accept such a result, and available reports do not confirm that the OpenAI claim has been recognized as a solution by the institute.
That earlier claim sharpened the issue now surrounding the larger collection. Tristan Buckmaster, a New York University mathematician working on Navier-Stokes, questioned whether people using AI systems may supply information that helps a model complete work already developed by researchers. He also warned that some results could take a researcher's partial contribution and carry it to completion. The concern is not that machine assistance is illegitimate; it is that attribution and originality become harder to assess when the system's prompts, intermediate work and training context are not fully available.
An independent advisory board hosted by the Institute for Advanced Study in Princeton, New Jersey, has recommended that AI companies release the prompts given to mathematical agents and the agents' chains of thought. The board said public release should begin rather than end the process of human understanding and incorporation into mathematics. Harvard mathematician Melanie Wood, a board member, said the aim was to establish standards that would let mathematicians understand AI-generated results and use them to advance the field.
OpenAI said it drew on the board's advice and public recommendations while preparing the release. Yet the board also opposed using advanced mathematical problems to test proprietary models. That position exposes a direct tension in the current evaluation model: difficult research questions can reveal what a system can do, but withholding the system and its operating record limits the community's ability to reproduce or challenge the claim.
The board has additionally called for equitable access to advanced AI models for mathematicians worldwide. Such access would not guarantee independent verification, but it would give more researchers the opportunity to examine whether a claimed capability is repeatable rather than a curated outcome from a private testing process. Similar concerns about reproducibility have shaped work at institutions such as MIT and Stanford, where computational methods are generally assessed alongside documentation of data, procedures and limitations.
The numerical scale is substantial: 377 results were released across 722 manuscripts and 372 result families; the model reportedly examined about 4,000 problems; the average attempt used about three hours of computing; and 10 solutions received explanatory summaries. The work spans five named mathematical areas and includes some unsolved problems. Those figures show the breadth of OpenAI's internal experiment, not a measured rate of reliable discovery, a comparison with human mathematicians or proof that the model can independently solve arbitrary research problems.
Lean checks are an important safeguard against certain logical errors, but formal verification has boundaries. The checker evaluates a formal proof written in its language; it does not settle whether the formalization captured the intended mathematical claim, whether the result is important or whether the model's route was original. Nor does the release report a public, repeatable evaluation in which outside mathematicians can run the same model under the same conditions.
The distinction resembles the difference between a validated instrument and a validated scientific conclusion. A checker can reliably test the object presented to it, just as a laboratory instrument can measure a defined quantity, but neither automatically establishes that the question was important, the experimental design was complete or the interpretation was original.
The central question is therefore not whether AI can produce impressive mathematical outputs. OpenAI's release provides evidence that a private system generated a large body of results and that many proofs could be checked mechanically. It does not yet resolve how much of the discovery was novel, how often the method works, how dependent it is on human-provided ideas or whether other researchers can reproduce the findings.
Mathematics has a stricter standard than a successful demonstration. Until prompts, computational records, relevant human contributions and reproducible access are available, these results should be treated as company-reported research claims rather than an independently established expansion of mathematical knowledge. The release is valuable precisely because it makes those standards impossible to ignore: AI may accelerate proof construction, but transparent attribution and human understanding remain the measures that determine whether acceleration becomes mathematics.
Formal verification means checking an argument against explicit logical rules rather than trusting a persuasive explanation. It can catch inconsistencies in the formal proof that reaches the checker, but it cannot reveal whether a model copied a known strategy, received a crucial hint or selected one successful attempt from many failures. In this case, that distinction separates a verified proof object from a verified claim of original mathematical insight.