The last high wall in the field of AI mathematics has collapsed, and GPT-6 Astra has fully conquered all tasks of FrontierMath Tier 4.
Just now, the last remaining FrontierMath Tier 4 hard problem has been solved by GPT-6 Astra.
At this point, every problem in this research-level test, once regarded as the "mathematics nightmare" for large models, has been successfully solved by AI at least once.
Epoch AI has also officially drawn its conclusion, FrontierMath Tier 4 is saturated.
Although as shown in the figure above, Astra has not directly reached 100%, and the score announced by OpenAI itself is 97.6%.
But according to Epoch's algorithm, "all solved" means that after accumulating the attempts of different models at different times, every problem in Tier 4 has had a successful solution.
And the one Astra solved is exactly the only problem that had not been cracked by AI before.
(Simply put, Astra might have made a mistake in a certain test, but the problem it got right was the one that had never been solved before.)
To be fair, this development speed is quite astonishing.
If you follow large model evaluation, you will know that on July 11, 2025, when Tier 4 was first launched, the highest score on the leaderboard was only around 5%.
Now just over a year and 2 months later, this once "high wall" has become a "stepping stone", and the last problem that was difficult for AI to bypass and solve through shortcuts has also been crossed.
Jay Pantone, the problem designer and associate professor of mathematics at Marquette University, also said that in the past, AI would always look for numerical shortcuts, but this time, the AI's solution is quite close to his own.
However, recently, when GPT-6 Astra has been focusing on making breakthroughs in the mathematics field and solving millennium-level problems every few days, Pantone said he can hardly be surprised anymore!
It can be said that mathematics has become one of the most prominent areas where Astra has demonstrated its capabilities recently.
From less than 2% accuracy, to the addition of a dedicated "research-level defense line"
FrontierMath was first released on November 7, 2024, and its original purpose was to prevent mathematical benchmarks from being quickly cleared by AI.
At that time, traditional mathematical benchmarks such as GSM8K and MATH had become increasingly difficult to widen the gap between top-tier models. Therefore, Epoch AI, together with more than 60 mathematicians, redesigned a batch of original problems that had never been made public before.
Among them are Fields Medal winners such as Terence Tao, Timothy Gowers, and Richard Borcherds.
After reading some of these research-level problems, Terence Tao said directly that these problems are "extremely difficult", and even judged that the hardest Tier 3 might keep AI stuck for several more years.
As expected, after the first round of testing, the leading model had an accuracy rate of less than 2%.
Specifically, at the very beginning, the core question bank of FrontierMath had a total of 300 problems, which were divided into three tiers according to difficulty: Tier 1, Tier 2 and Tier 3.
Tier 1 is roughly close to difficult undergraduate problems and mathematical Olympiad problems, but allows the use of more advanced tools; Tier 2 has reached the difficulty of senior graduate students; Tier 3 is closer to exploratory research problems that early-stage PhD students will encounter.
However, with the emergence of reasoning models, the first three tiers have gradually become insufficient to challenge AI, so in 2025, Epoch added a higher tier, Tier 4.
Most of Tier 4 problems are designed by mathematics professors and postdocs. Each designer conducts short-term research for about several weeks around their own research direction, and finally compresses the research results into a problem that can be automatically verified.
Initially, Tier 4 included a total of 50 problems, covering fields including analysis, number theory, combinatorics, topology, algebraic geometry...
At the time of release, after accumulating all historical test results of all models, only 3 problems had been solved, and these solutions also relied on some correct but insufficiently demonstrated assumptions.
In view of this, on the public sample problem page of the official website, Epoch once wrote: Some of these problems may not be solved by AI for decades.
(This sentence is now being constantly brought up as a funny past mistake.)
However, as models become more and more powerful, another problem has been exposed: the question bank itself can no longer withstand the "scrutiny" of AI.
OpenAI found during the test that there were more errors in FrontierMath than expected.
Subsequently, Epoch launched an independent audit, first using GPT-5.5 and Claude Opus 4.7 to screen for suspicious points, and then handing over the problems to mathematicians for review one by one. Finally, the v2 version was launched in June 2026, in which 12 problems in Tier 4 were corrected, 7 problems were removed, leaving 43 problems.
After the revision, the performance of the models has continued to rise all the way.
GPT-5.6 Sol achieved 83.0%, Claude Fable 5 reached 90.2%, and for GPT-6 Astra, the score has reached 97.6%.
More critically, Astra also solved the only remaining problem that had never been cracked by AI before.
At this point, every problem in Tier 4 has been successfully solved by AI at least once. This research-level defense line specially added for the most powerful models has finally been cleared.
However, this does not mean that "mathematics has been solved by AI". On the contrary, FrontierMath itself has moved to the next stage.
Now the entire project, in addition to Tiers 1-4, has added real Open Problems, as well as FrontierMath Erdős which formalizes Erdős open problems into Lean.
The former directly tests models with unsolved research problems that have not been resolved by the mathematics community; the latter requires AI to write complete proofs that can pass formal verification.
This time, Astra only solved 2 out of 68 problems in FrontierMath Erdős.
The story continues~
This article is from the WeChat official account "QbitAI", author: henry, published with authorization from 36Kr.