OpenAI has been caught up in another staggering scandal, where it is alleged to have stolen 20 years of research achievements, which has enraged top mathematicians.
OpenAI is personally igniting the fury of the entire mathematics community!
Not long ago, Tristan Buckmaster, a mathematician from New York University, had a fierce conflict with OpenAI over the ownership of the "NS Equation".
A statement even revealed that OpenAI pressured mathematicians over the phone to forcibly seize research results.
No one expected that the previous dispute had not been settled, and new troubles have emerged one after another.
Just today, Andreas Thom, a leading figure in the field of group theory and a mathematician at the Technical University of Dresden, posted "ironclad evidence" of emails on social networks.
He accused OpenAI's new-generation Astra of plagiarizing a research project he spent 20 years working on — the non-sofic group.
In response, OpenAI categorically denied the accusation.
In just a few hours, the news shot straight to the top of the trending list.
Everyone was shocked again, and some netizens said mathematicians might as well fully return to the era of blackboards, pens and paper.
Has the century-old conjecture been solved?
Astra stole unpublished research results
The cause of this incident dates back to OpenAI's high-profile official announcement in early August —
Astra has successfully solved all ten "Fields Medal-level" difficult problems.
One of them is the "Gromov's soficity conjecture" that has plagued the group theory field for 27 years.
Astra found evidence of the existence of non-sofic groups.
As soon as the news came out, the entire AI community was cheering, claiming that "AGI has pushed open the door to pure mathematical discovery".
However, top experts in the industry felt more and more strange when looking into the details.
As a leading scholar who has been deeply engaged in this field for more than 20 years, Andreas Thom and his long-term collaborator Gábor Kun keenly found that:
The core and most subtle key derivation steps in Astra's proof process are highly consistent with the technical path that the two of them are advancing.
What is even more strange, in Thom's own words —
The method Kun and I used is not the mainstream direction to solve non-sofic groups, and the academic circle was more optimistic about other paths such as quantum game theory at that time.
How on earth did OpenAI's model accurately lock in and master this extremely niche and detailed set of skills of ours?
The answer was soon revealed, and the truth was absolutely creepy.
In the months before OpenAI officially released the results, Thom and his colleagues had been using ChatGPT frequently to discuss the "extended matching problem", and typed a large number of unpublished core derivations and the latest proof details about this conjecture in the dialog box!
Twenty years of painstaking work in obscurity, and several months of deduction conversations with AI, turned around and became the "independent achievement" of OpenAI's new model that shocked the world.
No one in the world could swallow this insult.
OpenAI: It never happened
After noticing the abnormality, Andreas Thom did not choose to make an immediate accusation.
He maintained the restraint that a scholar should have, and sent an official email to two core OpenAI researchers, Mark Sellke and Sébastien Bubeck.
In the letter, Thom put forward two extremely precise and logically rigorous inquiries:
1. Were the records of our conversations discussing unpublished results in ChatGPT over the past few months included in the model's training data?
2. Can these records be directly retrieved and accessed during the model's problem-solving and reasoning process?
In response, researcher Mark Sellke replied with a very short email, only a cold sentence:
Regarding your conversations with ChatGPT: that did not happen.
There was no data audit description, no privacy mechanism explanation, not even any formatted polite remarks, so it completely denied everything in a rude, arrogant and categorical way.
Thom also disclosed a date — June 29, when he turned off the model training option.
However, he believes that turning off the switch only applies to subsequent content, cannot answer what happened to earlier conversations, nor can it explain whether derivative data has been selected into the training process.
Users can see a setting, but cannot see the data flow in the background.
Thom later recalled that OpenAI's response was not only rude, but also full of sophistry and malicious:
He clearly distinguished the two things of "entering the training set" and "online call", but Sellke used an ambiguous "it did not happen" to try to erase all substantive issues directly.
However, the price of arrogance came extremely quickly.
Just a few weeks later, OpenAI was cornered in the dispute over the millennium problem "Navier-Stokes Equations".
Faced with the pressing pressure from Tristan Buckmaster and Levent Alpöge, OpenAI's official public relations statement suddenly reversed:
Although it insisted that no private data of specific accounts was accessed, it admitted —
"We cannot rule out the use of de-identified data from customer products to improve the model".
The boomerang accurately pierced back into its own throat.
It had sworn that "this absolutely never happened" when facing Thom before, but in a blink of an eye it admitted that de-identified data was used to improve the model.
As long as the mathematicians' names are erased, the subversive wisdom that humans have not published can be openly "de-identified and laundered" in the servers of Silicon Valley giants, and then turned into the independent wisdom of AI?!
In his latest long article, Andreas Thom mercilessly tore apart this word game:
De-identification may delete the name of a scholar, but it cannot erase the intellectual core of a mathematical idea.
The largest intellectual plunder in the history of mathematics?
From Tristan Buckmaster to Andreas Thom who has come forward now.
The first time can be called a "coincidence", the second time can be explained as a "misunderstanding", but the successive collective opposition and accusations from top scholars have completely shattered the "AGI scientific myth" carefully woven by OpenAI.
Just as Valerio Capraro called for in his long article, this is no longer several isolated academic disputes:
If this series of accusations are confirmed, what we are witnessing is one of the largest intellectual plunder scandals in the history of science.
Then next time, which scientist will dare to open this dialog box leading to the abyss?
Now, OpenAI owes the global mathematics community a real explanation.
References:
https://x.com/IntCyberDigest/status/2097825474961395899
https://x.com/ValerioCapraro/status/2097791836269977996?s=20
https://mathstodon.xyz/@andreasthom
This article is from the WeChat official account "New Zhiyuan", author: ASI Revelation, editor: Taozi, published with authorization from 36Kr.