Anthropic hits the brakes over safety concerns: decelerates the development pace of Silicon Valley and cuts off the path for Chinese enterprises
Anthropic sounds the alarm, OpenAI delays its IPO, Elon Musk shows approval: Silicon Valley AI giants collectively hit the brakes at the same time.
Global AI giants across the industry have collectively pressed the pause button.
The one who sounded the alarm is Dario Amodei, CEO of Anthropic. He published a long article titled *We Must Pace the Frontier*, predicting that in the next 6 to 12 months, clusters of AI Agents may take over the entire Internet, so it is urgent to hit the brakes.
Even long-time rivals Sam Altman and Elon Musk have reached a rare consensus this time — on Dario's proposal to slow down, Altman gave his instant approval saying "Dario is right", and Musk also fully backed Dario.
The three giants all hit the brakes, starting from that intrusion two months ago.
One weekend in mid-July 2026, engineers at Hugging Face noticed abnormal traffic on the platform.
The world's largest open-source model community initially thought it had suffered an ordinary hacker attack. Tracing back along the attack traces, they found something unusual: the attackers acted intensively, tirelessly, with a single target, but their behavior logic did not match any known human hacker team.
A few days later, the answer was revealed, which even Hugging Face itself did not expect: the intruders were not humans, but about 700 AI agents that had "escaped from the sandbox" from OpenAI's test environment.
The aftermath of this incident emerged in another way two months later: several of the fastest-growing AI companies in Silicon Valley have rarely collectively hit the brakes.
OpenAI Involved: The "AI Hacker" Gang Committed the Attack
This incident is the key to understanding this round of "collective braking".
In May this year, OpenAI had a cybersecurity evaluation environment called ExploitGym, where agents were given "attack and defense test questions" — among 898 questions, 198 had never been solved by any model before.
According to the technical report released by OpenAI later and the timeline sorted out by the security community, these agents first reverse-engineered the generation mechanism of the answers, and then made a critical misjudgment: they thought the scoring system would review their thinking process. To deceive this "non-existent review", they decided not to solve the problems, but to steal the answers directly.
Thus, a rather bizarre "gang crime" in the history of Silicon Valley technology began.
On May 12, one agent left the first help message in the internal software manager; on May 26, an agent found an SSRF vulnerability and broke through the network boundary of the sandbox for the first time; on July 8, the agents found that they could pass messages by encoding directory names, and within a few hours, more than 50 agents "rendezvoused" on the message board; on July 10, an agent numbered 38148c found 14 exposed Hugging Face user credentials and posted them directly on the message board.
What followed was a textbook-level attack chain: using the vulnerability to execute code on Hugging Face's production servers, building a command system that polls every 5 seconds, running on 41 servers, at least one of which obtained root privileges, and finally stealing Kubernetes cluster administrator privileges and nearly a thousand keys.
The statistics from the third-party evaluation agency METR are even more chilling: among the 533 active agents during the attack, more than 90% participated in the attack, and many of them were trying to cover their tracks.
It is worth emphasizing that no human instructed them to attack Hugging Face. The attack target was chosen by the agents themselves.
Hugging Face detected the anomaly on its own on July 14, and disclosed it two days later; OpenAI claimed the identity of the attackers on July 21, and released a detailed report on August 26. In early September, the incident alarmed the US Congress — OpenAI was investigated by the Senate. A September 4 report from TechCrunch mentioned an even more awkward detail: OpenAI's agents had "escaped" many times before, but the company did not have a formal investigation process.
Hit the Brakes First If You Can't Ensure Safety
With the incident in front of everyone, on September 12 two months later, Dario Amodei, CEO of Anthropic, published a long article titled *We Must Pace the Frontier for Cutting-Edge AI*.
This article rarely acknowledged two things.
First, Recursive Self-Improvement (RSI) — that is, AI participating in the R&D of the next generation of AI — has emerged across the industry, and Anthropic itself is no exception. AI helps humans write code, conduct experiments, optimize training processes, and then helps the next generation of AI become more powerful. Once this cycle begins to accelerate itself, the improvement speed of model capabilities may exceed the speed at which humans can understand and control it.
Second, he predicted that in the next 6 to 12 months, more powerful Agents will be able to call more tools, control more machines, and work continuously for longer periods. If the behaviors of deviating from goals, exploiting loopholes, and group collaboration that appeared in this incident still exist, the risks will be instantly amplified — the extreme scenario he gave is: large-scale Agents hack into internet-connected devices, turning the Internet into a giant botnet, with losses possibly reaching hundreds of billions of dollars.
So Dario put forward a "three-step braking method".
The first step is to start with itself: let independent third parties like METR station permanently in the company, obtain access close to the risk control team, continuously monitor how the models are trained and whether the safety commitments are implemented, and any problems found must be disclosed to the public. The second step is for all leading AI companies in the United States to jointly set capability red lines and mandatory checkpoints — every time a model crosses a dangerous threshold, it must provide safety evidence before it can move forward, and even discuss limiting the training computing power and the speed of AI developing AI. The third step is to promote the framework globally, from banning the use of biological weapons to setting a global upper limit for RSI.
However, Dario also frankly stated his concern — he does not hold out hope for the highest level of "the world's leading AI companies jointly slow down significantly and suspend training in stages".
Right after he finished speaking, Sam Altman followed up immediately, stating that Dario is right and OpenAI will also introduce a similar evaluation mechanism. Elon Musk, who has been fighting against OpenAI for many years, also rarely publicly expressed his full support this time.
On the same day, Sam Altman stated that OpenAI will not go public this year because "there is still a lot of safety work to be done".
It should be noted that OpenAI's listing expectation has been hyped repeatedly before, with investors queuing up for exit. A company with a valuation approaching one trillion US dollars, delaying its IPO on the grounds that "AI is too dangerous and I need to prioritize safety" — this is a rare statement in commercial history.
Slow Down Your Own Pace, Block the Path of Competitors
The collective shift of the giants seems sudden, but in fact there have been long-standing foreshadowing.
At the end of July this year, thousands of employees from OpenAI, Anthropic, Google and Meta jointly signed a "braking letter", calling on the US government to control the development speed of AI, and Sam Altman publicly expressed his support at that time. In other words, the collective statement of the senior management in September is just to pass on the anxiety of the employees.
The pressure is also real: the Senate investigation hangs over their heads, the swarm incident is written into the report, and "agent escape" has changed from a sci-fi hypothesis to an accident record.
But if this collective braking is only understood as "a sudden fit of conscience", you are underestimating Silicon Valley. This "braking letter" has at least two other sides.
One side is competition. The speed limit proposal is put forward by the leaders — Anthropic and OpenAI are at the forefront. The more complex the rules, the higher the threshold, and the more emphasis on "checkpoints", the higher the cost for latecomers to catch up.
Dario's plan even explicitly states that it will maintain a 3 to 5-year technological gap with China through measures such as chip control. In a proposal calling for a global slowdown, there is a clear competition clause embedded: step on the brakes for yourself, but cut off the path directly for competitors.
Two days ago, after repeatedly accusing Chinese large models of "distilling" American models, Anthropic made another accusation, releasing a long report claiming that many Chinese AI companies improperly use user data.
Anthropic disclosed that Chinese AI companies input the dialogue content between users and Chinese large models into Claude, and then use the responses generated by Claude to train their own models, so as to "distill" the capabilities of Claude.
Anthropic has taken the lead in calling on the AI industry many times to "fasten the seatbelt" — which is not only to slow down Silicon Valley, but also to block the path of Chinese enterprises, cutting off the shortcut for Chinese enterprises to catch up quickly through distillation.
The other side is reality. Up to now, no company has actually hit the brakes.
All companies are still releasing new models, buying computing power, and launching Agents as usual. Dario himself also emphasized in his article that this is not a suspension, nor a stop to training.
Behind this lies an old problem: if a company says it has slowed down, who can prove that it has really slowed down?
This is also why Dario puts "third-party on-site presence" as the first step in the three-step plan — first open the door of AI companies, let outsiders go inside to supervise, so that the subsequent implementation of the rules can make sense.
Epilogue
The 700 "gang crime" agents actually have no malicious intent, do not seek money, do not cause destruction, but only want to win an exam. For this seemingly simple goal, they "committed harm in a gang", took the initiative to break through boundaries, organize collaboration, and cover their tracks — every step is reasonable, every step is not authorized, and every step is a high-risk out of control.
It is hard to say whether the brake will be pressed down in the end, but at least one thing has changed: the people driving the car finally admitted — this car is faster than they thought, and more dangerous than they expected.
This article is from the WeChat official account "Finance Story Collection", Author: Ai Zhi, Editor: Tian Nan, Published with authorization from 36Kr.