722 mathematical manuscripts, OpenAI has hidden 10 details worth understanding.
A single minus sign has affected three proofs.
At 10:00 AM Beijing Time on October 7, OpenAI published a batch of mathematical results generated by its internal model on GitHub. The first version of the catalog contained 722 manuscripts, categorized into 372 result families. In the catalog checked on October 8 in this article, the number of listed manuscripts dropped to 719, while the number of families remained 372.
The 3 removed manuscripts were retracted.
Among them, a geometric manuscript dated September 18 discussing the algebraicity of Weil classes noted in its retraction statement: "The proof contains a sign error". This means the proof has a sign error.
This proof requires that the positive and negative counts of geometric intersections are summed to zero before subsequent theorems can be invoked. The original expectation was that a certain operation would add positive numbers, but it actually added negative numbers. To use a simplified numerical example: the intended operation was "−2+2=0", but it was mistakenly executed as "−2−2=−4". The conditions for using the theorem were not met, and the two manuscripts that borrowed this construction also lost their supporting basis.
This batch of achievements has also received high praise.
On October 7, combinatorial mathematician Gil Kalai called this release "an amazing milestone for mathematics". Shortly after, he reminded that the proofs still need to be verified and digested by human mathematicians.
Discussions around AI mathematical research have long been underway.
In September, Claude's claimed proof for critical percolation sparked widespread discussion; on September 11, 25 Fields Medal winners including Terence Tao released a statement as initial signatories, acknowledging the progress of AI's mathematical capabilities, while criticizing the practice of enterprises that turn problem-solving into evaluation competitions and neglect understanding and knowledge inheritance.
On one side are important results that may change the entire field, and on the other side are retractions that have already occurred. What exactly did these manuscripts "break through"? Why can AI solve these difficult problems that once stumped mathematical geniuses in batches?
1
What Has AI Submitted to the Mathematics Community
This time, the large volume of mathematical materials released publicly by OpenAI in one go is extremely rich in content.
Open the public repository on GitHub, and you can directly access the paper PDFs and typesetting source files, allowing peers to check the arguments step by step. Some of the manuscripts are attached with Lean proof materials — which translate mathematical propositions and inferences into strict code for program verification. Interestingly, there are also 10 abridged inference abstracts that show part of the exploration process, including attempts that did not lead to a valid conclusion.
These manuscripts can also be called "preprints", which generally refer to research texts made public by authors in advance for others to read, discuss and check. Preprints can contain important achievements, but that does not mean they have undergone rigorous external review.
OpenAI stated that the vast majority of results come from the same workflow and the same internal model. If the important claims in these results are confirmed, it means that the capabilities of the same model have participated in research across multiple mathematical fields.
The announcement also converted the average computing power consumed by each result into the equivalent of ChatGPT Pro thinking for "roughly three hours".
2
722, 719, 372: Three Numbers, Three Different Meanings
722 and 719 count the number of manuscripts, while 372 counts the groups that gather related manuscripts together. OpenAI calls such groups "result families".
The repository description uses the term "related papers".
We can think of a family as a research folder. In this research folder, different manuscripts play their respective roles: some prove the main conclusion, some fill in a key step, and some generalize the conclusion to other conditions. At the same time, each has an independent manuscript, and they share one family number in the catalog.
In other words, one paper can contain multiple conclusions, and one folder may also answer different questions.
3
The Largest Family Has 14 Manuscripts in Total, What Is the Relationship Between Them?
Family No. 034 has 14 manuscripts, making it the largest family at present, belonging to algebra and complex geometry. The catalog includes both general arguments and versions specifically dealing with four-dimensional objects.
To understand why related research is divided into multiple manuscripts, we can first look at Family No. 088 which only has two manuscripts. The two study the same geometric indicator, roughly related to the "shadow" that a shape casts in all directions.
One of the studies focuses on what shape can minimize this indicator. The other provides a counterexample, proving that the originally guessed maximum value is not actually the maximum. They focus on the same indicator but answer different questions, so they are grouped into the same family and written as separate manuscripts.
This grouping phenomenon is very common. Among the 372 families, 205 families have only one manuscript, and 167 families have multiple manuscripts. Families with multiple manuscripts account for less than half of the total number of families, but contain 514 manuscripts, accounting for about 71% of the current total manuscripts.
4
Spanning 17 Fields, 60% of Families Are Concentrated in Seven Groups
According to the field classification in OpenAI's catalog, statistics show that 372 families are distributed across 17 fields. The seven largest groups have a total of 227 families, accounting for about 61% of all families. Theoretical computer science has 40 families, combinatorial mathematics has 37 families, and algebra and complex geometry has 36 families, ranking at the top of the list.
What are these fields respectively researching? For example, theoretical computer science studies how much computation algorithms require. Combinatorial mathematics mainly studies how discrete objects are arranged and counted. Algebra and complex geometry mainly studies the properties of shapes defined by equations.
Among the manuscripts released this time, there are 40 families and 73 manuscripts in theoretical computer science, while probability and statistical mechanics has 29 families and 105 manuscripts. The former has more families, while the latter has more manuscripts. If you only look at the number of papers and the grouping according to related achievements, you will get different statistical results.
5
"Breakthroughs" Have At Least Five Forms
What exactly does the "breakthrough" of the 722 manuscripts announced by OpenAI this time mean?
Based on the manuscripts we sampled and read, we can divide the contributions into five types. This is not a hierarchical classification given by OpenAI, and one paper may cover several of these types at the same time.
(1) Provide an answer to an unknown problem. For example, π can be continuously approximated by fractions. The research here focuses on: how fast can the error decrease infinitely many times relative to the denominator? The new manuscript claims that the "irrationality measure" indicator that measures this speed is equal to 2.
(2) Make the conclusion applicable to more scenarios. When studying particles and electromagnetic fields, many existing results require the initial state to be small enough or have special symmetry. The new manuscript claims that under the model specified in the paper and other initial state conditions, these two types of restrictions can be removed, while still ensuring that the solution remains smooth for a long time.
(3) Cross a boundary that was once guessed to be impossible to break through. For example, the longer an integer is, the more calculation steps multiplication usually requires. The new manuscript claims that under the same fixed machine model, the growth of the number of steps can be slower than the long-guessed limit.
(4) Fill in the prerequisite that many conclusions rely on. For example, many studies assume that the Unique Games Conjecture is true before proving that certain algorithm tasks cannot reach a certain approximation accuracy. If the new manuscript holds, these conclusions no longer need to list it as an unproven assumption.
(5) Provide another proof for a known answer. For example, a geometric counterexample manuscript first acknowledges: "Their result already gives a negative answer", which means that previous results have already given a negative answer, and then proposes a direct construction. What is added here is a new proof route, and it remains to be seen whether it can bring clearer explanations or methods that can be used by other research.
6
Seven Highlights: First Look at Which Statement They Changed
What exactly is the "breakthrough" of this public release?
For a specific achievement, we can compare three things together: what previous researchers have already proved, where the new manuscript claims to push the boundary, and what conditions it still retains.
The seven cases below are presented in this order.
Take multiplication as an example: the paper compares how fast the number of calculation steps grows as the number of digits of integers continues to increase. The new manuscript claims to reduce this growth rate to a lower level.
If it holds, it will advance the theoretical research on computational efficiency. At the same time, whether real devices can be accelerated as a result depends on how the algorithm is implemented and how long the integers being calculated are.
It is worth noting that the "Quasi-Riemann Hypothesis" and "CM Hodge" in the table still need to be interpreted with their scope clearly defined. The former still does not reach the full Riemann Hypothesis, and the latter is limited to CM abelian varieties.
7
Which Three Manuscripts Were Retracted and What Impact Did They Cause
The three retracted manuscripts are "Algebraicity of Weil classes on split abelian eightfolds", "Algebraicity of Kuga–Satake Correspondences for K3 Surfaces", and "The rational Hodge conjecture for products of K3 Surfaces".
The latter two both involve a type of surface called K3, and both use the adapted version of the flawed construction in the first manuscript.
However, the retraction statement does not declare the original proposition to be false: the gap in this argument does not mean that the proposition has been overthrown, nor does it rule out the existence of another valid proof.
The three retracted manuscripts all belong to Family No. 032. After the retraction, this family was reduced from 8 manuscripts to 5, and the main Hodge manuscript on CM abelian varieties is still retained. CM abelian varieties are a type of geometric object with special arithmetic symmetry. This remaining main manuscript is also connected to a research published 27 years ago.
8
A New Result That May Fill the "If" Left 27 Years Ago
In a paper published by mathematician James Milne in 1999, he proved a connection: As long as one conjecture holds, the other conjecture will naturally hold. What was missing at that time was the proof of the former conjecture itself.
Specifically, the former is the rational Hodge conjecture for all CM abelian varieties over complex numbers, and the latter is the Tate conjecture for all abelian varieties over finite fields. The former problem is limited to a class of geometric objects with special arithmetic symmetry, while the latter problem enters an operation system that only has a finite number of elements.
This manuscript claims to have filled in the proof of the former conjecture. If it holds, it can connect to the proven connection in Milne's paper and deduce the latter conjecture. The impact of the new result will therefore go far beyond the objects it directly studies.
On October 7, Milne added a note in his paper catalog mentioning OpenAI's announcement, using the phrase "it seems" to describe this consequence. This shows that he has also noticed the connection between the new manuscript and the old theorem.
9
Why These Problems May Be Suitable for AI
OpenAI stated that the starting point for expanding mathematical research testing was "evaluations saturated", which means that the original test questions are increasingly difficult to distinguish the capabilities of the model. OpenAI has provided the model with a total of about 4000 research problems, and then screened and grouped the outputs.
These problems have one thing in common: even if there is no known answer to the problem, the research task can be written very clearly. A public task lists the research object, input conditions, the conclusion that needs to be proved, and what kind of counterexample is sufficient to overturn it.
The proof route needs to be found by the model.
The public abridged inference abstract records an attempt to design a test. This test needs to make it easy for answers that meet the requirements to pass, while blocking answers that do not meet the requirements. However, the model also found that one scheme exposed a structure that could be exploited, allowing answers that did not meet the requirements to slip through.
Adding more dense random perturbations can prevent this kind of fraud, but it will also reduce the probability that answers that should pass can pass the test. The model then continues to search: how to plug the loopholes without breaking other guarantees.
As a result, "looking for a proof" has become more concrete, rather than just an "adjective" that describes an ongoing effort. It can be split into repeatable steps: propose a method, check where it will fail, then adjust the route, and of course call on theorems that have been proved by predecessors in the process.
Looking back, several favorable conditions can be seen from these materials: The goal is clearly defined, existing research provides callable tools, and part of the results can also be subject to formal verification.