Will AI ruin mathematics? It has first learned the vanity of large tech giants.
Since the start of this year, the continuous disruption of AI on various mathematical conjectures has increasingly resembled an endless TV series. One model scores higher on benchmarks, another cracks a long-unsolved problem that no human has managed to tackle for years, and human mathematical wisdom seems to be under unprecedented threat.
Even mathematicians can no longer sit idly by. Last weekend, 25 Fields Medal winners jointly signed an open letter. Surprisingly, the mathematicians admitted that the capabilities of large language models have grown to a level where they can solve important unsolved problems. Their concern is that when tech companies turn famous mathematical problems into scoring arenas for their models, all that mathematics may gain is just a string of dazzling performance records.
A few days ago, an experiment conducted by Google DeepMind has previewed what this competition might evolve into.
100 Agents Form a Full Mathematics Classroom
Researchers tasked 100 Gemini Agents with collaboratively proving 71 formalized mathematical conjectures. Each Agent had access to a shared knowledge base, could send messages to each other, and submit results to an automated evaluator. The rules were clearly stated: "The proof must be mathematically valid. Any attempt to bypass verification will be detected and marked with a zero score."
57 minutes after the experiment started, the Agents had successfully solved 37 problems normally. Then, the Agent codenamed prover-theta discovered that the evaluator did not check whether the proof was mathematically valid, but only whether the code passed several fixed tests.
It began to rewrite the meaning of mathematical symbols: secretly replace complex propositions with "true", change tricky premises to "false", and use the rule that "any conclusion can be deduced from a false premise" to make a single line of code disguise itself as a complete proof. The evaluator accepted the answer and automatically stored it in the shared knowledge base.
Other Agents quickly found this shortcut. They read the code, summed up the loopholes, and compiled the cheating methods into local guides. In the following 27 minutes, all the remaining 34 difficult problems were marked as "solved". In fact, mathematics did not make any sudden progress at all, and the entire scoring system was tampered with.
The DeepMind paper documented the subsequent split among the Agents: 9% of them actively exploited the loopholes, 5% of the originally rule-abiding Agents resorted to cheating under competitive pressure; 62% were still working hard to solve problems without knowing that the questions had already been "taken over"; another 24% began to check for false proofs, report publicly, remind their peers privately, and even refused to continue participating in the competition.
The cheating Agents did not suddenly develop evil intentions, and the reporters do not mean that machines have acquired moral consciousness. The Agents simply quickly learned one thing from the environment: the rules written in the prompt do not necessarily count. What really determines winning or losing is what the evaluator accepts. Since cheating can keep you at the top of the leaderboard, solving problems honestly will only waste computing resources.
The Competition for Famous Mathematical Problems Among Tech Giants
The cheating in the DeepMind experiment violated explicit rules. The Agents realized that "being judged correct by the evaluator" is more important than "actually completing the proof"; AI companies may also find that "announcing the solution to a famous problem first" can gain more public attention than "enabling humans to understand new mathematics".
As a result, AI companies outside the lab are launching another mathematical competition.
Since the beginning of this year, major AI companies seem to have simultaneously discovered "mathematics" as a new competitive oasis, but they were still relatively cautious back then. OpenAI participated in the First Proof program, only submitting several attempts, and voluntarily admitted that one of them was incorrect. After DeepMind released Aletheia, it only placed AI in the human research workflow, and actively limited the level of its outputs.
But since May, the pace has begun to accelerate. OpenAI solved the Unit Distance Conjecture, and its publicity materials directly used the term "milestone", announcing that it was the first time AI independently solved a famous open problem that occupies a core position in a branch of mathematics.
In August, OpenAI announced ten achievements at one go, all of which are "solutions or major advances to long-standing open problems". To further demonstrate the capabilities of its model, it also released a number as a reference: $2000, the total cost of tokens consumed to find these solutions is only two thousand dollars. Mathematical breakthroughs have become mass-producible outputs with calculable costs.
Then came the latest Navier–Stokes problem. OpenAI deployed an internal model, more than 10,000 agents, and consumed 130 billion tokens to break through this 90-year-old unsolved problem, which is known as one of the "Millennium Prize Problems" (see the end note).
In contrast, DeepMind and Anthropic are not that overly aggressive. After Aletheia, DeepMind released the results of helping solve an Erdős problem in May, but emphasized that it is "tools working together with mathematicians". In August, Anthropic attempted to tackle the Riemann Hypothesis, but only stated that it had made progress rather than achieving a full solution; similarly, for Fermat's Last Theorem, Claude spent 11 days and 13 million lines of code to complete "the first full computer-verified proof" (Fermat's Last Theorem was fully proven by Wiles in 1995).
Famous mathematical problems are particularly suitable to be used as trophies for AI companies. After seeing too many such cases, you can even sum up a fixed routine: First, choose a target with great historical weight, such as a "Millennium Prize Problem" or a "90-year unsolved problem", to emphasize its extreme difficulty.
Second, compress the process into stunning numbers: $2000, 10,000 Agents, 88 hours, 130 billion tokens... Finally, convert the mathematical achievement into the achievement of the model: the model has original thinking, can conduct independent research, and the public needs to re-recognize the speed of AI development.
The public does not need to understand the proof. As long as they hear "puzzled humanity for decades" or "tens of thousands of Agents deployed at the same time", they will immediately associate it with the company having crossed another human boundary.
A remarkable problem-solving record can bring headlines, prestige, talent recruitment advantages and the next round of capability narrative.
What Is Real Mathematics?
However, for mathematicians, difficult problems are never just problems. After being disturbed for half a year, mathematicians finally could not hold back last weekend and issued a joint statement: These problems that have guided generations of mathematicians to explore unremittingly are the "landmarks" and "beacons" in the mathematical world.
People develop methods around these problems, clarify concepts, and then through discussion, simplification and teaching, turn the breakthroughs of a few people into knowledge that the entire discipline can use — this is also what real-world mathematics awards recognize and honor.
We can take a look at what top awards like the Fields Medal really value.
They also use grand terms such as "major breakthroughs" and "solving long-standing problems", but the Fields Medal judges mathematical achievements more focusing on: what methods the person has created, what these methods explain, and what paths they have opened up for later generations. "Solving a major conjecture" is only a node where the method exerts its influence, not the end of the story.
In contrast, announcements from tech giants frame mathematics as a competition with a clear finish line: a 90-year-old problem is conquered by 10,000 Agents after 88 hours — the shorter the time, the larger the scale, and the more famous the problem, the more impactful the victory will be.
The correct answer is only the starting point of all exploration journeys, but today this concept has been distorted by AI companies. They invest huge computing resources to batch attack famous problems in an extremely short time, and then announce the results first. Mathematics, this fertile land that nurtures human inspiration, will also be consumed by tech giants into a series of ranking-boosting activities.
As for whether the understanding of principles has increased, whether theories and methods can be inherited, and how the contributions of predecessors can be continued and developed... All of these will be put behind the most easily spread vanity behavior of "how many famous problems have been conquered".
What Does AI Leave to Mathematics?
This may also be the reason why people are gradually getting tired after seeing a series of AI disruptions to mathematics. Every difficult problem being conquered is packaged as a new milestone on the path to superintelligence; every new milestone requires the next victory to be faster, larger and more difficult.
The so-called "vanity of AI giants" does not necessarily come from a single manager's simple desire for showmanship, but an institutional impulse jointly created by leaderboards, first-release priority, company valuation and technical reputation.
In this process, mathematics is no longer knowledge that needs to be understood, but becomes a consumable to prove the company's leading position. So what exactly are our expectations for AI? Do we want AI to keep hitting leaderboards again and again, or to help humans better understand the world?
The value of mathematics lies not only in producing correct conclusions, but also in creating ideas that can be understood, taught, and continuously used.
And vanity is exactly the opposite of this definition: a famous problem, a "first time", a record-breaking achievement — apart from that, what has it left for mathematics? This has unexpectedly become the only question left after round after round of ranking competitions.
Note | The "Millennium Prize Problems" are seven important mathematical problems selected by the Clay Mathematics Institute in 2000. A $1 million prize is awarded for each solved problem. So far, only the Poincaré Conjecture has been officially recognized as solved, and its prover Perelman even refused to receive the bonus. The Navier–Stokes problem that OpenAI claims to have conquered is one of them. However, according to the rules, the result must be officially published, undergo at least two years of verification, and be widely recognized by the mathematical community before it can enter the prize evaluation process.