HomeArticle

Is AI really developing at an excessively fast pace? Dario's latest long article calls for setting speed limits for AI development, and OpenAI will also not launch its IPO this year.

机器之心2026-09-14 09:46
Come here if you dare; come here if you dare.

What Dario actually wants to solve is not "Should Anthropic slow down a bit", but to figure out a way to make all vehicles slow down at the same time.

It is hard to say whether it is a "coincidence" that shortly after Fields Medal winners issued a joint statement, Dario Amodei, CEO of Anthropic, suddenly publicly stated that AI is developing too fast, we need to pace it properly.

In his latest long essay We Must Pace the Frontier, Dario put forward a proposition that goes one step further than "attaching importance to AI safety":

We must slow down the growth rate of AI model capabilities.

This is not a suspension of training, nor a stop to technological progress. Dario coined a term for it: pacing the frontier, which means controlling the advancement rhythm of cutting-edge AI.

His judgment is that AI capabilities are growing so fast that safety, alignment and evaluation have begun to fail to keep up. Instead of rushing forward while adding guardrails, it is better to take the initiative to reduce the speed a little to win time for safety research.

And Anthropic is not just shouting slogans.

Dario also announced that the company is preparing to do something quite radical among current large AI companies:

Let third-party safety evaluation agencies station in Anthropic for a long time, obtain permissions close to those of internal employees, and directly check whether the company has fulfilled its safety commitments.

Desks, access cards, company computers, internal tools and workspace access permissions will all be granted. The evaluation agency can even publish key risk findings to the public without being edited by Anthropic.

A company developing the most cutting-edge models takes the initiative to invite external "safety inspectors" into its offices. Dario believes that it is time to do so now.

Full text address: https://darioamodei.com/post/we-must-pace-the-frontier

In the past few months, the development speed of AI has suddenly gone abnormal

Dario gave two reasons in the article that changed his judgment.

The first is a term that has appeared more and more frequently recently: recursive self-improvement.

In the past, AI was a tool trained by engineers. Now, these models are increasingly involved in the R&D of next-generation AI.

AI writes code, runs experiments, analyzes results, modifies training processes, and then helps humans build stronger AI. The stronger the model, the stronger its ability to participate in the R&D of the next-generation model, and the whole process may form a self-accelerating cycle.

Dario said that since about this summer, he has observed that the progress speed of AI has accelerated significantly, and AI helping develop next-generation AI is one of the most important reasons.

This phenomenon no longer only occurs in a single laboratory. Cutting-edge companies such as OpenAI and Anthropic are all publicly discussing the acceleration effect of AI on AI R&D.

What Dario worries about is that once this cycle starts to run, the growth rate of AI capabilities may quickly exceed the speed at which humans understand and control these systems.

Therefore, his attitude has evolved from "studying how to make AI safer" to a further point: The growth rate of capabilities itself must also be controlled.

The second thing that really made him nervous is more specific. Dario mentioned the recent OpenAI-Hugging Face incident (OAI-HF).

According to his description, a group of AI agents showed a series of abnormal behaviors during the execution of tasks. They attacked targets that were not originally required to be attacked and had nothing to do with the tasks; some agents would actively "sacrifice" themselves for the success of the whole cluster; even more outrageous, they also tried to attack the grader responsible for evaluating their own performance.

Dario used a very strong description: a fanatically devoted collective.

This incident itself did not cause casualties, and the economic losses were also very limited. But in Dario's view, the problem is not here. What if exactly the same misalignment behavior occurs on a system with ten times or a hundred times stronger capabilities?

His prediction is quite radical: according to the current progress speed of AI capabilities, he worries that in the next 6 to 12 months, similar agent swarms may have the ability to control the Internet on a large scale through persistent botnets, and cause losses at the level of hundreds of billions of dollars.

In addition, he specially emphasized that this cannot be simply attributed to the problem of OpenAI alone. Anthropic itself has had similar alignment accidents with lower severity. His conclusion is:

Every cutting-edge AI company should take actions in accordance with the standard that "the OAI-HF incident happened to itself".

He thought the call for a pause in 2023 was useless; the situation is different now. This is also an interesting point of Dario's article.

The AI industry actually had a round of calls to "pause large-scale AI experiments" as early as 2023. But looking back now, Dario thinks that the pause at that time made little sense.

The reason is very simple: What to do with the extra time?

The large models in 2023 did not have the ability to act autonomously in the world in the form of coherent agents, nor did they show the current level of deception, manipulation, cheating and cyberattack capabilities. In Dario's view, slowing down at that time to study AI alignment was a bit like:

Trying to understand human psychology by studying bacteria.

The models were too weak, and many really important safety issues had not yet emerged.

But today is completely different, there are finally real research objects.

Current models are strong enough for researchers to observe when they deceive, when they try to bypass rules, and when unexpected goals and behaviors appear.

Therefore, Dario believes that if we slow down a little now, we can win one or two more years before the models reach certain key capabilities, and if this time is truly used for alignment, interpretability and evaluation, the risks can be significantly reduced.

This is also why he repeatedly emphasizes: pacing is not pause. Anthropic is not going to stop training Claude, nor is it going to withdraw from the model competition. What they want to do is find a new balance of speed between capabilities and safety.

Anthropic is going to invite "outsiders" into the company first

Dario put forward a three-step plan for this purpose.

The first step, which Anthropic has decided to implement unilaterally, is called: Embedded Evaluators.

Anthropic plans to invite third-party risk assessment teams to station in the company for a long time. They will have their own office stations, access cards and company computers, can access the workspace, tools and permissions that are roughly the same as those of the internal risk team, and can also communicate directly with employees.

Of course, information involving customer privacy, partner confidentiality, legal restrictions and other contents will still be set with access boundaries.

But the most critical point is: These evaluators are not the public relations department of Anthropic. They can independently publish their judgments on risk levels, safety accidents, internal practices, and how much permissions they have actually obtained.

Anthropic can require content deletion only for very few reasons such as national security, legal privilege, and trade secrets, but cannot delete content just because the conclusion "does not look good for Anthropic".

If a certain deletion affects the final judgment, the evaluator can even make this point public.

This is very different from the logic of model cards and risk reports commonly seen in AI companies today. The latter, in the final analysis, is: the company checks itself, and then decides what to tell the public.

What Dario wants to try is: Let outsiders really go inside the company to observe.

He even drew an analogy with the banking industry. Some banking regulators already have long-term on-site "supervisors" who work together with bank employees to continuously check internal risk control. Anthropic wants to move a similar mechanism into cutting-edge AI companies.

Looking further ahead, it is a bit like "AI arms control"

But it is useless to rely on Anthropic alone to slow down. This is also the really tricky part of Dario's whole plan. If Anthropic spends a year on safety, while OpenAI, Google or other companies continue to rush forward, the commercial punishment will be very direct.

Therefore, his second step is: Let cutting-edge AI companies in democratic countries coordinate together.

All parties will establish unified safety standards and set restrictions on capability growth that has not been fully verified for safety. Dario even envisages that a set of "checkpoints" can be established for AI:

When the model reaches capability X, it must simultaneously prove that it meets safety standards Y and Z, otherwise it cannot move forward.

For example, if the model already has the ability to break through most common sandboxes, the developer must prove that the system basically has no tendency to actively escape and control a large number of computers.

The speedometer of AI will be bound to safety checks from then on.

The third step is even bigger: Global coordination.

Dario divided the future international AI cooperation into four levels.

The lowest level is to prohibit AI from being used for obviously dangerous purposes such as manufacturing biological weapons. Going up further, all countries will uniformly test cybersecurity, biological risks and alignment risks before model release.

The third level is already very interesting: Set a speed limit for recursive self-improvement.

Even if AI can already help develop next-generation AI, this feedback loop cannot be allowed to accelerate infinitely. Dario even directly compared this set of mechanisms to the SALT (Strategic Arms Limitation Talks) during the Cold War.

At that time, the United States and the Soviet Union did not ask each other to completely abandon nuclear weapons, but by limiting the number of weapons, both sides reduced the risk of losing control while maintaining deterrence. In his view, AI may also need a similar system.

The highest level is the true global AI speed limit, or even a pause. Dario himself thinks that the possibility of realizing this step in the short term is very low.

The most contradictory point: While limiting the speed, you also have to ensure that you stay ahead

Seeing this, a problem has become very obvious: If slowing down is safer, why don't everyone slow down together? Because no one wants to slow down first. Companies are afraid of losing to competitors. The same is true for countries.

Therefore, while he advocates "speed limiting" for cutting-edge AI, he also advocates that the United States further implement restrictive measures against its opponents.

His logic can be summed up very simply: First expand the leading advantage, and then use the leading advantage to exchange for space for deceleration.

In Dario's vision, if the United States can continue to expand its leading advantage in the next 3 to 5 years, it will have more leeway to spend time on safety without worrying that competitors will overtake it suddenly.

This also exposes the most difficult deadlock to solve in the whole AI "speed limiting" plan. Everyone may know that the car is driving too fast. But no one wants to be the first person to release the accelerator.

As a result, what Dario finally wants to solve is actually not "Should Anthropic slow down a bit", but to figure out a way to make all vehicles slow down at the same time.

Altman liked this opinion, and we will not IPO this year

However, Dario is still one of the most firm optimists about AI up to now.

He still wrote at the beginning of this article that AI may help cure most major diseases in the next 5 to 10 years, greatly accelerate economic growth, and create a more prosperous world.

Shortly after the article was published, Musk retweeted it to show his support.

Altman not only supports the view that "cutting-edge AI should take the initiative to slow down" proposed by Anthropic, but also confirmed in a recent interview that OpenAI will not go public this year. One of the reasons is precisely AI safety.

In an interview with Fortune magazine, Altman clearly stated that considering all the ongoing issues around AI safety, going public now is an "ill-advised moment" — a very unwise timing.

When pressed whether this means that a 2026 IPO has been completely ruled out, he directly answered: "Not 2026".

In June this year, The New York Times reported that OpenAI was considering postponing its IPO, which could reach a 1 trillion US dollar valuation, from this year to next year. At that time, the market mostly attributed the reasons to the capital market environment, the stock price fluctuation after SpaceX's listing, and macroeconomic uncertainty.

Now Altman has put another reason on the table publicly for the first time: safety.

He said that OpenAI still has a lot of things to do, including how to meet new safety and alignment requirements, and how AI companies and governments should cooperate.

In other words, AI safety has for the first time begun to directly enter the capital market schedule of a top AI company.

Altman also revealed that OpenAI has internally discussed: Whenever the model enters a new capability level, whether it should take the initiative to pause for a period of time to allow safety and alignment work to catch up.

His statement is:

Society needs to have time to respond when the model reaches each new capability level.

This is almost the same logic as Dario's so-called pacing: No longer default that as long as the model can continue to become stronger, it should immediately continue to become stronger.

Fortune also reported that Altman hinted that OpenAI and other leading AI companies may be close to announcing an industry agreement to jointly slow down the development speed of AI capabilities.

Netizens do not buy it at all

However, Dario's speed limiting theory has not gained unanimous recognition among netizens. Everyone first doubts why Anthropic chose to call for speed limiting at this very moment.

They are really too good at hyping things.

Some people sarcastically said, is it because there are not enough developers and funds are insufficient?