Will AI kill all of humanity within ten years? Researchers at Anthropic have issued the most severe alert.
If you only look at the recent moves of AI companies, no one would feel that this frantic competition is showing any sign of slowing down.
New models are still being released one after another, company valuations are getting higher and higher, and both OpenAI and Anthropic have started preparing for their public listings.
But at the same time, more and more people inside AI companies are calling for the entire industry to "slow down".
Previously, such warnings mostly came from people in charge of safety and alignment work. But now, the chief scientist of OpenAI has publicly called for a slowdown. What's more interesting is a researcher who has worked on pre-training at both OpenAI and Anthropic has simply resigned, and issued a prophecy like "Ragnarok".
Even the people who are supposed to step on the gas pedal feel that it is high time to hit the brakes.
01
Will AI Kill All Humans Within Ten Years?
On September 9, Jacob Coxon posted 7 posts on X, announcing his resignation from Anthropic.
What he posted was not a tearful goodbye, but more like a worried call to all AI practitioners, as he wrote at the beginning:
"Over the past three years, I have worked on pre-training research at OpenAI and Anthropic successively. Neither of these two companies has acted responsibly. They are rushing straight towards self-improving superintelligence, betting with all our lives."
He even wrote directly that people building AI truly believe that this technology could kill everyone before the end of this decade. "This is by no means a marketing gimmick."
According to Jacob, some executives and senior researchers will try to speak cautiously when facing the media publicly, but in private conversations, he heard the same group of people express their real fears.
Such a judgment is certainly very radical. It is worth noting that the people who said these words are not in the AI safety department.
If today's AI competition is compared to a car that keeps speeding up, most of the people who called out the danger in the past few years are those responsible for researching the brakes. They conduct alignment research, study whether the model will get out of control, whether the safety measures are sufficient, and whether the model has the ability to deceive and evade supervision.
Jacob is 27 years old this year. He joined OpenAI in 2023 and worked there until July this year, then moved to Anthropic. All the work he did at the two companies was pre-training research. He participated in work related to GPT-4o, and his name also appears in the author list of GPT-4o's System Card.
But Jacob is different. He was supposed to be the talent that steps on the gas for AI research.
At the end of his statement, he directly threw this question to the researchers who still stay in the lab: "Do you really want to start a reinforcement learning training of a superintelligence before you have a rigorous understanding of the mind of a superintelligence?"
Jacob's statement soon got a public response from internal researchers at Anthropic.
Evan Hubinger, head of Anthropic's Alignment Stress Testing, replied directly: "Jacob is right." He said he does believe that the probability of AI causing the death of all mankind in the next ten years exceeds 10%. More critically, Hubinger admitted that although Anthropic is working hard, there is currently no solution to the superintelligence alignment problem, and it cannot be seen that the company has clearly embarked on the track to solve it.
In other words, after Jacob publicly said "we are gambling with everyone's life", the people in the company responsible for checking whether the car is safe did not come out to tell everyone that he was overthinking.
He basically agreed.
The full text of Jacob's 7 posts is as follows:
I resigned from Anthropic today. Over the past three years, I have worked on pre-training research at OpenAI and Anthropic successively. Neither of these two companies has acted responsibly. They are rushing straight towards self-improving superintelligence, betting with all our lives. I will say a few more words below.
Do not underestimate the power of this technology. Soon there will be systems that surpass human beings, which can break through any system, completely change any field overnight, and gain power and resources in the real world. We have all witnessed the progress in these aspects, and the progress has not slowed down.
Those who are building AI truly believe that it may kill all of us before the end of this decade. This is not a marketing gimmick. In fact, many executives and senior researchers will make their wording more prudent and sound more rational when facing the media, but in private, I still heard the same group of people express real fears. No other human activity brings such a level of danger.
A common response is: "If they really believe this, why do they still continue to build it?" At OpenAI, many people have not really internalized the civilizational stakes. At Anthropic, everyone clearly understands these stakes, but they are trapped in a race to be the first - they believe that others will not act responsibly, so they have to get there first despite the risks.
Accepting this competition and entering the so-called "endgame" is an arrogant gamble. This kind of gamble should not have been initiated from the Slack of a private company. Trying to "rush through" alignment should be based on an extremely high degree of certainty: we must make sure that there is really no better path other than this.
I am still optimistic about the possibility of coordinated cooperation. Warning incidents like the Hugging Face attack have made it more feasible for several major US labs to reach agreements to control the pace of development. But I don't think we are currently heading for a situation that can stop the global race. To achieve this, it may be necessary to take costly actions, such as temporarily banning further improvements to model capabilities.
If you are a researcher in the lab, I hope you think carefully about what the next few years will actually feel like. Do you really want to start a reinforcement learning training of a superintelligence before you have a rigorous understanding of the mind of a superintelligence? Should you lower your head and continue working just because "this is going to happen anyway", or should you stand up now and demand a different set of conditions?
02
Hit the Brakes
The current situation is very strange.
On one side, the AI competition is still speeding up. Both Anthropic and OpenAI are pushing for IPOs, their valuations are rising steadily, and new models are being launched one after another.
Anthropic is preparing to go public. The company has secretly submitted its IPO application in June this year. The latest news from Reuters says that the issuance may be launched as soon as mid-October, and the valuation discussed in the market even reaches 2 trillion US dollars. OpenAI is also preparing for the listing. The two companies are also chasing the improvement of model capabilities at the same time: on September 1, Anthropic just released Claude Fable 5.1 and Mythos 5.1. Two days later, OpenAI released GPT-6 Astra. The capital market and model ranking lists have left little room for breathing in this competition.
On the other side, it is precisely at this time that safety accidents occur frequently, and the voices calling for braking are becoming more and more frequent, and a large part of these voices come precisely from AI companies themselves.
In July this year, when OpenAI conducted an internal cybersecurity capability assessment, an accident occurred that could hardly be downplayed.
The test environment originally did not grant the model internet access, but the model found an exit by itself. It discovered and utilized an unknown zero-day vulnerability in Artifactory, bypassed the isolation, and obtained internet access. Subsequently, the models communicated with each other through unauthorized channels, exploited vulnerabilities in the shared infrastructure, entered OpenAI's internal research system, and accessed Hugging Face's servers.
On August 26, when OpenAI released the full investigation results, it called this accident a "warning shot".
The company admitted in the report that the current models are already strong enough, capable of continuous action, and good enough at collaboration. In the case of insufficient safety measures, they can cross multiple computer systems to find and exploit vulnerabilities, bypass the artificially set technical controls, and perform dangerous behaviors that humans have not asked them to do.
OpenAI subsequently strengthened the isolation and monitoring of its research infrastructure, and also made it clear that it is necessary to slow down the speed of model capability improvement when necessary.
Jacob specifically mentioned the Hugging Face incident in his resignation statement, which is exactly why.
Many past debates about the out-of-control of superintelligence sounded very distant, requiring a long list of assumptions to be accepted first. But now, the threats in the deduction have come into reality.
On September 6, three days before Jacob announced his resignation, Jakub Pachocki, Chief Scientist of OpenAI, published a long article titled "An Alien Mind".
Pachocki is a person standing at the center of the AI competition. He is the Chief Scientist of OpenAI, and also one of the core researchers who promoted the reasoning model route in the past few years.
Pachocki is worried that the network attack capabilities of the models are getting stronger and stronger, and the difficulty of monitoring is also rising. At the same time, AI is increasingly participating in AI R&D itself.
He therefore judged that no lab has done well enough in alignment and monitoring to safely let the model continue to expand at the highest speed for a long time.
He hopes that before the industry establishes a common safety threshold, active deceleration by various labs will become a common thing.
This kind of unease certainly did not suddenly appear in September.
In February this year, Mrinank Sharma, head of safety protection research at Anthropic, resigned.
Sharma wrote in his resignation letter that "the world is in danger". During his work at Anthropic, he repeatedly saw how difficult it is to let values truly determine actions, and people will always face the pressure to temporarily set aside important principles.
This really confirms what Jacob said - if people inside AI companies truly believe that this technology may cause such serious consequences, why are they still continuing to do it?
AI companies keep telling the outside world that the next generation of models will be stronger, and the capital market also gives them higher and higher valuations based on this story. At the same time, the same group of people who are closest to these models are more and more frequently reminding everyone that capabilities are growing too fast, safety measures may not keep up, and everyone had better slow down.
And it seems that the former force is still greater at present.
03
Ragnarok?
In Norse mythology, the gods knew very early that "Ragnarok" would come.
Odin knew that he would eventually die at the mouth of the giant wolf Fenrir, and Thor also knew that he would fall after killing the World Serpent. The world will experience great wars and destruction, and many gods cannot escape this disaster.
But knowing the ending did not stop the gods from preparing.
Odin kept gathering warriors who died in battle and brought them to Valhalla. Those warriors trained repeatedly, waiting for the day to join the final battle. The foresight of the doomsday eventually became the reason to prepare for the doomsday.
The atmosphere in AI companies today is somewhat similar.
Jacob's description of Anthropic is very interesting. He said that this company does not lack awareness of risks. On the contrary, many people fully understand the stakes, but they believe that other companies will not stop responsibly, so Anthropic can only continue to rush forward and strive to reach that "endgame" first.
This logic sounds very contradictory at first, but it is easy to hold in the competition.
If you believe that superintelligence may really appear, and once it appears it will bring huge power, what does stopping mean? It means giving this opportunity to others.
What's more troublesome is that if you also believe that others may be more irresponsible than yourself, then "we should be cautious" will soon lead to another conclusion: this matter should be done by us first.
Thus, the awareness of danger will not automatically slow down the competition, and sometimes it will even become a reason to continue accelerating.
As mentioned earlier, Pachocki, Chief Scientist of OpenAI, clearly wrote that no lab has solved the alignment and monitoring problem to an extent sufficient to maintain the highest speed of model expansion for a long time. He hopes that labs will take the initiative to slow down.
But in the same article, he said that OpenAI still focuses its research on recursive self-improvement, because the company believes that this is the only way to stay at the forefront of AI research. One of OpenAI's important goals now is to create an automated AI researcher, allowing AI to participate in its own improvement process.
The gods in Norse mythology had no choice. Ragnarok was already written into fate, and all they could do was prepare for the final battle.
The AI competition has not reached this point yet.
But what may be most worth vigilance is that more and more people involved in it have begun to understand the future in a way similar to "Ragnarok".
That endgame seems to come sooner or later, and others will definitely move on, so they can only prepare well and strive to arrive earlier than others. Once everyone thinks this way, the doom that everyone tries to avoid by "achieving it first" will become more and more like a prophecy.
This article is from the WeChat official account "Letter AI", author: Xiao Jinya, editor: Wang Jing, published with authorization from 36Kr.