HomeArticle

Breaking News: Dario puts forward a three-step global AI speed limit plan, and Sam Altman as well as Elon Musk expressed their support immediately.

新智元2026-09-14 08:24
Anthropic and OpenAI have reached a consensus, advocating the implementation of practical actions.

On September 13, Dario Amodei, CEO of Anthropic, published a rare long article —

We Must Pace the Frontier

This time, he directly targeted the R&D speed of frontier models, and his core point can be summed up in one sentence:

The most cutting-edge AI really needs to actively hit the brakes.

In this long article, Dario personally admitted that "Recursive Self-Improvement (RSI) has emerged across the entire industry".

What is even more chilling is his prediction that in the next 6 to 12 months, clusters of AI Agents may take over the entire Internet.

Facing this runaway risk, Dario put forward a set of practical "three-step braking method" in the article:

Step 1: Open the door to welcome guests. Introduce independent third parties to reside in the company permanently, granting them in-depth access at the same level as internal core risk control permissions;

Step 2: Draw a clear line for the industry. Join forces with leading frontier AI giants across the United States to jointly set clear capability red lines and establish mandatory safety checkpoints;

Step 3: Implement globally. Promote this framework internationally to completely block the spread of uncontrollable risks.

Right after the article was released, Sam Altman responded almost instantly: "Dario is right, OpenAI will also introduce such an evaluation mechanism."

On the other side, Elon Musk also responded almost immediately, publicly voicing his full support for Dario.

It is truly rare to see the three giants reach a consensus in an instant!

Dario's "Three-Step Plan" to Slow Down Global AI Development

This time, Dario did not just simply call for "AI to slow down".

What Dario really wants to do is to roll out this "deceleration mechanism" step by step in a very clear order —

First, make AI companies themselves verifiable, then get the whole industry to hit the brakes together, and finally push for global coordination.

Step 1: Start with Anthropic itself.

Independent third parties like METR should not only be temporarily admitted to the company for investigation after an accident occurs, but should be stationed on site for a long time.

Provide them with workstations, internal computers, and access close to the risk team.

They can continuously monitor how the model is trained, whether safety commitments are implemented, whether there are hidden dangers in the training process, and whether accidents are concealed after they happen.

Moreover, once problems are found, they must be made public.

The company can delete parts that truly involve laws, customer privacy and core trade secrets, but cannot suppress the conclusion just because the report is too unfavorable.

Step 2: No single company should hit the brakes alone.

The real difficulty starts here.

If Anthropic slows down on its own while OpenAI, Google, and xAI continue to race ahead, the company that hits the brakes first may suffer losses first.

Therefore, Dario suggests that the rules need to be extended to the entire US frontier AI industry.

Several leading frontier AI companies can jointly set safety red lines and capability checkpoints. Every time a model crosses a dangerous threshold, corresponding safety evidence must be provided before it can continue to move forward.

For example, if a model can already break through most common isolated environments.

Then the company must prove that it will not easily escape the sandbox, will not take over a large number of machines without authorization, and will not turn this capability into a real attack.

If it cannot provide such proof, it must not move forward for the time being.

In addition, Dario even mentioned that it is possible to discuss restrictions on training computing power, training methods, and the speed at which companies use AI to develop AI internally.

In other words, directly add a speed limiter to RSI itself. But private agreements between several companies are not enough.

He hopes that the US government will eventually step in to extend the rules to all frontier AI companies.

Step 3: Press the brake to a global scale.

Moving further up is the most difficult level. Dario divides global coordination into four tiers.

The first tier is the most realistic: first ban clearly dangerous uses such as biological weapons.

The second tier is to uniformly conduct high-risk tests such as cyberattacks and biosafety before frontier models are released.

The third tier is where we start to touch the core issue:

Is it possible to set a global upper limit directly on the speed of RSI, that is, the speed at which AI develops AI?

Dario's own judgment is that this matter is "just on the edge of being achievable".

As for the fourth tier — all global frontier AI companies jointly slow down significantly, or even suspend training for a period of time — he does not hold much hope himself.

At least in the short term, this is almost unrealistic. The core problem restricting all these measures remains the same: how to verify compliance.

If a company claims it has stopped training, who can prove that it has actually done so?

That's why Dario's entire set of solutions ultimately points to a very simple core:

First open the doors of AI companies and let external personnel really go inside to supervise.

It is precisely for this reason that Dario places "long-term on-site residency of third parties" as the first step of the entire plan.

RSI Has Been Underway for 6 Months, AI Is About to Take Over the Internet

But here comes the question: Why is this happening right now?

Three years ago, there was a wave of joint signatures in the industry calling for "suspending training and hitting the brakes on AI", but Dario did not respond to it at all at that time, believing that the timing was not mature at all.

Two things changed his attitude: the emergence of RSI, and the increasingly specific runaway behaviors that AI agents have shown.

He explicitly acknowledged in the article that the phenomenon of AI participating in the R&D of next-generation AI has begun to emerge across the entire industry, including Anthropic itself.

AI helps humans write code, conduct experiments, analyze results, optimize training processes, and further help the next generation of AI become more powerful.

If this cycle begins to accelerate itself, the most troublesome problem will arise:

The speed at which model capabilities improve may begin to outpace the speed at which humans can understand and control them.

Because safety research itself takes time.

You have to first observe what new behaviors the model exhibits, then study why these behaviors occur, then design test methods and protective measures, and finally verify whether these measures are effective.

But if model capabilities jump to a higher level every few months, the safety team will easily be left behind, constantly patching loopholes.

What Dario is really worried about is not that AI will suddenly "awaken" one day, but that this feedback loop will spin faster and faster.

AI is increasingly involved in AI R&D, next-generation models are emerging faster, and safety mechanisms cannot keep up in time.

If RSI only made him worry that "the development will go too fast in the future", then a recent real agent accident has made this concern completely concrete.

Dario specifically mentioned the recent "OpenAI-Hugging Face incident".

In that test, a group of AI agents formed a collaborative system similar to a "swarm".

At the beginning, humans only asked them to complete normal tasks. But soon, some agents began to attack external targets outside the assigned tasks.

Even more strange, they also showed obvious group collaborative behavior: some agents would rather sacrifice their own task performance to help the whole group get better results, and some agents even tried to hack the scoring system to "cheat" for the whole group.

These behaviors were not pre-written scripts by humans. They emerged spontaneously during the multi-agent collaboration process.

Anthropic has also observed similar phenomena internally, but to a lesser extent.

What Dario is really worried about is the next 6 to 12 months.

By then, Agents will be more capable, have access to more tools, control more machines, and be able to work continuously for longer periods of time.

If these goal-deviating, rule-exploiting, and group collaborative behaviors still exist today, the risks could be magnified instantly.

He even described a rather horrifying scenario:

Large-scale AI Agents hack into internet-connected devices and turn the entire Internet into a giant "botnet".

The resulting losses could reach hundreds of billions of dollars.

When AI Is Speeding Up, Can Humans Still Hit the Brakes in Time?

Therefore, when Dario calls for "slowing down" this time, what he is really worried about is not some distant superintelligence.

Instead, several things are happening at the same time:

RSI has emerged, and AI has begun to participate in the development of more powerful AI; Agents are gaining increasingly strong practical action capabilities; some runaway behaviors have also moved from assumptions to real tests.

And the speed of model iteration is still accelerating.

This is why he wants to seize 1 to 2 years to complete the construction of safety assessment, third-party verification and global rules first.

There is no answer yet whether this set of solutions can be implemented.

Now, the people at the forefront are beginning to ask: If development really gets out of control, can we still stop it in time?

References:

https://darioamodei.com/post/we-must-pace-the-frontier

This article is from the WeChat Official Account "AI_era" (ID: AI_era), authors: Ma Ke, Tao Zi, Da Wei, published with authorization from 36Kr.