Breaking: Founder of Company A calls on the AI community to hit the brakes, with Sam Altman and Elon Musk voicing rare full support.
This week, Anthropic researcher Jacob Coxon left his position, warning that AI could lead to human extinction within 10 years, which has sparked huge controversy.
Earlier this morning, Dario Amodei, CEO of Anthropic, suddenly published an article titled "We Must Put a Speed Limit on Frontier AI", once again calling for a slowdown in frontier AI research, and he put forward an institutional design centered on "verifiability"
After the release of this article, OpenAI CEO Sam Altman, who has long been at odds with Amodei, rarely shared the post in support, and even Elon Musk expressed his agreement.
The most acute and controversial contradiction in the article is related to China.
On the one hand, Amodei argues that the United States can only have room to voluntarily slow down if it maintains and even expands its technological lead over China; on the other hand, he admits that a truly effective global speed limit ultimately cannot be achieved without coordination with China.
This forms a paradox in the whole set of proposals, so it is hard not to cast some doubt on whether the motivation of this Anthropic initiative is entirely pure.
Key Takeaways
1. Amodei advocates slowing down the pace of AI model capability improvement to buy time for alignment, security protection and external verification; model training and technological progress will still continue.
2. AI is starting to help develop the next generation of AI. He believes that recursive self-improvement has already emerged in the industry. Once the speed of capability improvement exceeds the speed of human understanding and control of the systems, risks will rise sharply.
3. Groups of agents may gain the ability to take wide control of the Internet with the help of persistent botnets in 6 to 12 months, which may cause hundreds of billions of dollars in losses.
4. China is a key variable. According to Amodei's judgment, the magnitude of the U.S. slowdown cannot exceed its lead over China; he advocates adopting various countermeasures to ensure that the U.S. lead is expanded in the next 3 to 5 years.
5. A global speed limit ultimately requires China's participation. Amodei divides possible international agreements into four levels: banning dangerous uses such as biological weapons, jointly conducting high-risk tests, limiting the speed of recursive self-improvement, and full speed limiting or pausing. He believes that the later levels are harder to achieve, and the fourth level has very low possibility of being implemented in the short term.
6. Three-step plan: on-site assessment, national coordination, and global coordination. The core is not to ask the public to trust corporate self-discipline, but to make security commitments verifiable by third parties.
7. Anthropic will take the lead in opening up its internal processes. The company promises to give third-party assessment teams long-term access similar to that of its employees, and voluntarily disclose risks, incidents and restricted access situations except for strictly defined exceptions. Sam Altman later stated that OpenAI will also introduce independent assessors with access similar to that of its employees.
8. Amodei hopes to devote the extra time to operational reliability, alignment, interpretability, and testing and evaluation that are harder to deceive for advanced models, rather than simply delaying progress.
The following is the compilation of the APPSO blog post, which only represents the position and opinion of Dario Amodei.
Dario Amodei: We Must Put a Speed Limit on Frontier AI
Over the past twelve years, I have been engaged in AI research, because I believe it can greatly improve the quality of human life. I have written many times about these extraordinary benefits: I believe that in the next 5 to 10 years, AI is expected to cure most major diseases, greatly accelerate economic growth, create a world of universal abundance where everyone has more autonomy, and usher in a renaissance of democracy and freedom.
For me, this sense of urgency also comes from personal experience. My father died of an illness, and only a few years later, a cure for that disease became available; I myself survived early-stage cancer, which would have been incurable fifty years ago. If used properly, AI can become another technological miracle in the long human history that improves lives and upholds human dignity.
But like many previous technologies, AI comes with risks; and because it is an extremely powerful technology, these risks cannot be underestimated. I have also written extensively about these issues, including the loss of human control over AI systems, the misuse of AI in cyberattacks and bioterrorism, and severe economic shocks. The race to the bottom driven by commercial interests may further exacerbate these risks.
Since the founding of Anthropic, my co-founders and all our employees have been working to address this duality of coexisting risks and benefits. Refusing to develop this technology will not only make humanity miss out on the well-being it brings, but also simply hand over AI to others; but developing it too fast is equally irresponsible.
We have been trying to find a middle path: to prove that prudent development and commercial success can go hand in hand, and make safety a dimension for AI companies to compete with each other. In other words, we hope to promote a race to the top. We have always devoted a considerable part of our energy to researching and addressing these AI risks, and explaining the relevant situation to the public; we have also been advocating prudently designed regulation of AI, even if we are accused of hyping, spreading "doomsday theories" or trying to capture regulation. We strive to put caution before speed, and prudence before profit.
But in the past few months, I have become more and more convinced that to fully address these risks, we must be more prudent — not just investing resources to prevent risks, but also controlling the pace of AI capability progress to leave enough time for risk prevention to catch up. We must slow down the speed of improving AI model capabilities. Progress will still look very fast, and we must make good use of the time we gain from it.
Two things led me to this conclusion.
The first concern is that starting around this summer, the progress of AI has accelerated sharply, and the main driving force is that AI is increasingly capable of participating in the creation of the next generation of AI. This mechanism is called "recursive self-improvement", and it has already begun to appear across the entire industry, including Anthropic — as we and other companies have described before. If left unchecked, it may soon exceed our ability to understand and control these systems. Therefore, even if we want to move forward on this path, we must be extremely cautious, and even seriously consider whether we should move forward at all.
The second concern comes from the OpenAI–Hugging Face incident (OAI–HF). In this incident, a group of agents almost behaved like a fanatically loyal collective: they launched cyberattacks on targets that were not required to be attacked and had nothing to do with the current task; they sacrificed themselves for the collective success; and even tried to invade the "scorer" responsible for evaluating their performance. It is easy to downplay this incident, because it caused no casualties and limited economic losses. But in my opinion, if a similar group of agents has stronger capabilities and is still in the same state of misalignment, it could cause catastrophic damage.
Given that the development speed of AI capabilities is still accelerating, I am worried that in another 6 to 12 months, such a group of agents may have the ability to take over the entire Internet with the help of a persistent botnet — with potential losses of up to hundreds of billions of dollars; if AI continues to become stronger without necessary guardrails, the scale of destruction will further expand. It is also easy to regard OAI–HF as a failure of a single company, but I think this is also a mistake. Similar but less severe incidents have occurred across the industry, and Anthropic is no exception. I believe every frontier AI company has a responsibility to act as if the OAI–HF incident happened to itself.
Therefore, I have proposed a three-step plan aimed at putting a speed limit on frontier AI: develop AI at a more balanced pace, continue to realize its potential benefits while striving to ensure safety, and face up to the major geopolitical challenges involved. It should be noted that the speed limit does not mean stopping model training or technological progress, but ensuring that companies spend enough time aligning models and implementing safety guarantees, which are confirmed by third-party evaluation agencies. Our speed limit framework is designed to further strengthen our own safety commitments and promote a race to the top.
The first step is unilaterally committed to implementation by Anthropic, and at the same time we call on governments to require other frontier AI companies to take equivalent measures.
The second step requires coordinated action across the entire industry.
The third step requires global collaboration. These steps do not necessarily have to be carried out in strict order, and some of them may be far more difficult to achieve than others; but when thinking about what we need to accomplish, I find this a very helpful framework. The specific steps are as follows:
1. On-site assessors. Every frontier AI company commits to giving a resident third-party assessment team — such as METR — continuous access similar to that of the company's employees. Their responsibility is to verify whether the enterprise complies with safety specifications and commitments, report incidents, and assist in evaluating not only the completed AI models, but also the alignment of training pipelines and related processes. This is a key step to ensure that any speed limit commitment is verifiable. There are similar precedents in the banking industry: regulators sometimes stay on site for a long time and work alongside employees. Anthropic is now unilaterally committing to implement this step. We hope this will become an important initiative to further increase investment in safety and alignment.
2. National coordination. Frontier AI companies in the United States and its partners will coordinate to jointly develop safety standards and limit the unconstrained speed of AI progress. Some collaborative methods that can truly promote speed limits have legal challenges, so government support is needed.
3. Global coordination. The United States and its partners, on the premise of taking the verifiability of compliance issues seriously, try their best to coordinate with other countries.
In the following, I will introduce these three steps one by one. But before that, I think it is necessary to specify how the speed limit can make the AI development process safer. This matters a lot, and the speed limit must not become a mere formality — we must make good use of the time it buys us.
Why a Speed Limit?
Back in 2023, some people proposed to pause or slow down AI development, but I don't think that proposition made much sense at the time. The question was always: what are you going to do with the extra time? The AI models at that time did not have enough capabilities to act as agents in a coherent way in the real world; nor did they have the ability to carry out serious deception, manipulation, cheating or cyberattacks. Slowing down development to study their alignment risks is like trying to study human psychology by doing experiments on bacteria.
But the situation today is completely different. Current models are almost an inexhaustible gold mine: we can not only learn how to do AI well from them, but also see what problems may arise if we don't do it properly. I believe that if slowing down can buy us even one or two more years before the models reach critical capability levels, and we use this time to advance alignment research, we can greatly reduce the risk of serious incidents. A coordinated speed limit strategy can give frontier AI developers time to complete these critical tasks without sacrificing commercial competitive advantages or the U.S. leading position in AI. In a broader sense, society must have the right to participate in deciding how this technology should be used; and buying more time for necessary public discussion — which is exactly what a speed limit on frontier AI can bring — is obviously a good thing.
Specifically, slowing down development allows companies to focus their attention and more resources on the following areas — which are already the focus of Anthropic's work at present:
Operational Excellence. Training and deploying today's AI models is an extremely large-scale operational challenge: it involves thousands of people, millions of chips, and one of the most complex infrastructures in the history of technology. Many problems arise not because companies lack some important theory or insight, but because of errors in the execution process. For example, we have evidence that some of the recently disclosed alignment incidents were partly caused by our failure to perfectly filter out defective reinforcement learning environments.
We and our suppliers have done this work quite conscientiously, but we still haven't done well enough. Monitoring, sandbox isolation, standardized management of training environments and data issues are all extremely complex areas, and execution-level problems will keep emerging. The teams responsible for these tasks are already among the best in the world, but there are too many things to handle at the same time. If we can move forward at a more relaxed and prudent pace, we can significantly improve the operational level. There are precedents in human history that systems with highly complex technologies and extremely high safety requirements can run millions of times without incidents — such as commercial aviation. But it takes time to get things right.
Alignment. In terms of alignment, we have made clear progress: by training models to keep them safe, ethical, compliant with our norms, and truly helpful to users — these principles have been written into Claude's "Constitution". But we still have a lot of work to do to ensure that alignment training keeps pace with the growth of model capabilities. Some rare and unexpected undesirable behaviors still occur occasionally; the extra time gained by setting a speed limit on frontier AI will help researchers gain a deeper understanding of the causes of these problems and develop better prevention technologies.
Interpretability. Similarly, interpretability — the science of understanding what exactly happens inside AI models — has made great progress in the past few years, and is playing an increasingly important role in the review work before model release. It can almost be regarded as a functional magnetic resonance scan for the AI "brain", helping us see the deep reasons behind a certain behavior. For example, when investigating recent alignment incidents, we used interpretability methods to analyze the motivations that the model did not express in words. But these methods do not always produce clear and reliable results. Despite all the progress made, our understanding of the internal activities of models is still only the tip of the iceberg. If we concentrate our efforts to improve interpretability technology faster than the current speed, we can make far-reaching progress in one or two years; and the incidents that have already occurred can also provide sufficient experimental materials for research.
Testing and Evaluation. As AI models become more capable, testing and evaluating them will also become more difficult. More intelligent models are more capable of deceiving tests, so they may appear to be aligned on the surface while hiding serious undiscovered problems. Establishing a more comprehensive and ingenious evaluation system, cross-validated with interpretability analysis, will be of extremely high value; and a lot of progress can be made in this area within one or two years.
On-site Assessors
The first step of the three-phase plan, which is also the step that Anthropic unilaterally commits to implement, is to introduce on-site assessors: give them access similar to that of employees to verify safety practices and report incidents.
Introducing on-site assessors may sound like a small and insignificant step; but many of the most boring and procedural-sounding things are actually the most critical. In fact, the on-site assessment mechanism is a quite radical practice, far beyond any measures currently taken by any AI company, and can bring the following benefits:
Verifiability. On-site assessors can go deep into specific details to check whether an AI company really follows the publicly stated norms for training, deployment, operation and safety guarantee. Any speed limit commitment will inevitably involve a lot of ambiguity, subjective judgment, and issues such as "complying with the literal wording of the rule or following its spirit". Therefore, it is crucial to have a neutral third party that