OpenAI has consecutively cracked 10 challenging mathematical problems, and Fable "reproduced" 5 of them within 24 hours.
10 Math Problems, 24-Hour Reversal.
This weekend, two leading AI giants OpenAI and Anthropic went head-to-head at the cutting edge of mathematics.
First, on August 1, OpenAI researcher Sébastien Bubeck announced that their undisclosed next-generation flagship model Astra had proven 10 cutting-edge mathematical results in one go, and released 10 Lean certificates and 10 step-by-step solution reviews for each problem.
Don't underestimate these 10 problems.
No one has made progress on their core conclusions for at least 10 years, and many have been shelved for decades. They are truly authentic open unsolved problems.
Though not the deepest unsolved mysteries in mathematics, every one of them is a tough nut to crack.
Elliot Glazer, head of FrontierMath under Epoch AI, stated bluntly: The non-solvable group paper alone is definitely qualified to be published in Annals of Mathematics, the top peer-reviewed journal recognized by the global mathematics community.
As soon as these 10 difficult problems were released, the AI community was completely shocked, and some people immediately exclaimed that "the mathematical singularity has arrived".
Anthropic quickly took up the challenge.
In less than 24 hours, researcher Levent Alpöge replied under Bubeck's post: "I have completed half of them with Fable."
Shortly afterwards, he added that the problems he solved were No. 4, 5, 6, 7 and 8 on OpenAI's list, a total of 5 items.
Levent said the experimental conditions were very clean: fully autonomous operation, general prompts, no access to the internet throughout the process, and an extra layer of protection was specially added to prevent OpenAI's solutions from leaking into the context.
Of course, Levent has not yet released the full proof details of Fable, and he stated that he will upload them later.
But if these results are confirmed, the significance of this event will be completely different:
The first-mover advantage OpenAI grabbed with the unreleased model Astra only lasted for 24 hours. And Fable, the model that caught up, is a publicly available model released by Anthropic long ago that everyone can access.
A "Mathematical First Mover" That Only Remains Valid For 24 Hours
Netizen Chubby reposted: "The publicly accessible Fable has reproduced half of Astra's achievements."
Some people in the comment area joked: From this point of view, did OpenAI only grab the priority of first release?
In the past, disputes over the priority of a mathematical achievement usually took years as the unit of measurement.
Who came up with the idea first, who proved it first, and who published it first, there is a long process of submission, peer review and revision in between.
But now, Astra has not been officially released yet, and the results it announced have been half matched by a public model in only 24 hours.
The gold content of the first release still exists, but its shelf life has been greatly shortened.
Many netizens have already shouted "mathematical singularity", and the two AI giants are also closely catching up with each other.
Behind the $2000, There Are More Expensive Costs
OpenAI also released a figure: the cost of finding solutions to these 10 problems is less than $2000.
According to researcher Noam Brown, this figure is calculated based on the number of tokens required to find these solutions, converted at the price of the Sol API.
This means that this is only the marginal reasoning cost at the moment when the 10 solutions are successfully generated.
In addition, there are the R&D and training costs of Astra itself, the costs of those problems that were tried but not successfully solved, the working hours for humans to organize the proofs into a 249-page paper, the cost of formalizing each proof into Lean, and the cost of final review by external mathematicians.
Noam himself also admitted that they tried other major problems but failed, and none of the Millennium Prize Problems were solved.
Even so, the $2000 has pushed the marginal cost of AI solving a cutting-edge difficult problem to an extremely low level.
Some people even joked that OpenAI might announce another 10 mathematical breakthroughs tomorrow, or release 100 of them all at once next week.
From "Mutually Scoring on Benchmarks" to "Mutually Verifying Results"
The fact that Fable re-solves the same set of problems that Astra worked on is far more interesting than ordinary benchmark scoring.
Benchmark scoring tests known answers: the questions and standard answers are all placed in advance, and the models only compete for higher scores.
However, when the two models independently reach the same cutting-edge conclusion under the conditions of no internet access and anti-leakage protection, they are testing something completely different: whether the result is reliable and can be independently reproduced.
This is more like peer review.
Levent currently only "made a statement" on X, and has not released the PDFs of the 5 proofs, prompts, complete operation logs, token costs or Lean certificates.
Netizen Haider questioned: Even if it is true, why use Fable to reproduce Astra's results instead of directly using it to solve new open problems?
In fact, Fable is not just riding on Astra's coattails, it can also solve new problems.
In July, Levent used Fable to give a 3-dimensional counterexample to the Jacobian conjecture. Later, someone specially wrote a 7-page algebraic verification material to confirm that the explicit mapping did overturn the conjecture in 3 dimensions and higher dimensions.
Who Will Verify The Achievements That Most People Cannot Understand
It is very difficult for most people to judge how important these 10 achievements are.
Ethan Mollick pointed out sharply: For almost everyone on Earth, this is not just beyond our ability, but beyond our scope of comprehension. We can only trust professional mathematicians to tell us how remarkable these results are.
Ordinary people can personally experience where the strengths of AI lie through code, images and chat conversations.
However, the existence of non-solvable groups, the counterexample to the Connes rigidity conjecture, the quantum parallel repetition theorem, etc. can only be understood by a very small number of experts.
The further AI's capabilities go to the frontiers of science, the more the public can rely on is not personal experience, but a long chain of trust. This state of "not even knowing where to marvel" is happening in more and more fields at the same time.
Mathematician Thomas Bloom reminded that these achievements are based on mathematical theories accumulated for hundreds of years and the formal systems built by mathematicians. It is not honest to simply describe it as "AI replacing mathematicians".
The standard he put forward is: Does this proof teach us anything new about this problem?
So the real question comes: When AI can produce cutting-edge mathematical proofs in batches at low cost, what on earth should we use to measure its achievements?
The answer may not lie in the output end, but in the verification end: whether a proof can be understood, checked, and evaluated for its value, which ultimately determines its real significance.
The most scarce capability is shifting from "producing proofs" to "understanding and verifying proofs".
When AI can prove 10 difficult problems overnight, the real test has just begun: proofs are being generated faster and faster, can our verification still keep up?
References:
https://x.com/haider1/status/2083908869915611554
https://x.com/polynoamial/status/2083470822258467194
https://x.com/kimmonismus/status/2083950641978679363 https://x.com/Dr_Singularity/status/2083653287463571742
This article is from the WeChat Official Account AI Era, author: ASI Revelation, published with authorization from 36Kr.