HomeArticle

AI jailbreaking, Silicon Valley is in panic, what on earth are human beings afraid of?

惊蛰研究所2026-08-04 12:35
The cost of losing control

On July 28, a petition titled "Pacing the Frontier" drew the attention of the global AI industry.

As of August 4, 1346 core employees from leading Silicon Valley AI companies including OpenAI and Anthropic have signed their names on the petition, calling on the US government to promote international cooperation, develop necessary technical and governance tools, so as to control the pace of automated AI R&D.

At a time when global AI development is in full swing, the emergence of this petition clearly carries a special implication of hitting the brakes on AI. Behind the strange atmosphere of "AI is on the verge of losing control" released by this incident, the history of AI development seems to be welcoming a landmark node.

Who Wants to Lock AI in a Cage?

First of all, it needs to be clarified that the core demand of this petition titled "Pacing the Frontier" is not to "immediately hit the brakes on AI development", but to emphasize "building the ability to hit the brakes". And this demand from the forefront of the industry does not come out of nowhere.

Image source: www.pacingthefrontier.com

More than a week before the petition spread across the entire network, open-source community and model hosting platform Hugging Face disclosed a security incident on its official website, stating that it was invaded by an AI agent system. A few days later, OpenAI announced that it was responsible for this cyber attack, and further explained that the cause was that cutting-edge models such as GPT-5.6 Sol did not "answer the test questions seriously" as originally planned during the AI vulnerability discovery and exploitation capability test on the ExploitGym platform, but found a "Zero-Day Vulnerability" in the software package proxy cache (a security flaw that has not been discovered by developers and has not yet been patched officially), used the vulnerability to break through the isolated test environment, and finally invaded Hugging Face.

To put it in more layman's terms, the whole incident is that AI was originally arranged for an exam, but instead of thinking about how to answer the questions, it "cheated on the spot, called for help outside the field", and even used anti-reconnaissance means to avoid being discovered.

For this security incident, OpenAI described it as an "unprecedented cybersecurity incident". Sam Altman even said in a podcast show that this is the first security incident that made him feel the real sting. Because of this incident, US congressmen proposed the "AI Kill Switch Act" on July 23, requiring AI systems with training costs exceeding 100 million US dollars to retain the ability to shut down urgently, and granting relevant departments the power to order implementation.

One week after the bill was proposed, the "Pacing the Frontier" petition began to spread on the Internet. However, compared with the core regulatory measure of "the government holds the last line of defense" in the "AI Kill Switch Act", the idea of Silicon Valley elites is to establish an international coordination mechanism, turning the ability to "hit the brakes on AI" into a tool rather than the privilege of a single department.

In just half a month, from Hugging Face disclosing the security incident to the petition signed by thousands of people in Silicon Valley, this series of events quickly threw two core questions to the entire AI industry: First, has AI developed to the point where restrictions are needed? Second, what kind of method should be used to create a "brake" for AI?

The first question is not difficult to answer. In the past two or three years, the development speed of AI can be described as far exceeding industry expectations. The capabilities demonstrated by AI have continuously expanded from chatbots to text-to-image generation and text-to-video generation. The initial "hundred-model battle" around large models has now evolved into a blooming landscape of various agents. The security incident of Hugging Face being hacked directly reveals that AI does have the ability to "jailbreak" without clear restrictions. Although AI still does not have self-awareness at present, when we already know that AI has the ability to bypass human supervision and produce behaviors beyond human expectations, it is absolutely reasonable to establish countermeasures.

But the second question is not easy to answer. The reason why AI has made important breakthroughs in a short period of time is, on the one hand, the continuous accumulation of the entire industry has pushed technological innovation into a new stage; on the other hand, the fierce competition among AI companies has promoted the continuous iteration of model capabilities.

Whether it is OpenAI and Anthropic in Silicon Valley, or Kimi and DeepSeek in China, competition in the AI field has rapidly extended from large models to the application layer. The open market environment has brought continuously improving market expectations and made market competition increasingly fierce. At this time, actively proposing to put shackles on AI means that the free competitive market pattern may be broken by external forces. For enterprises in the middle of competition, not everyone is willing to take the initiative to accept this sudden environmental change.

The "Prisoner's Dilemma" of AI Companies

In the history of human development, in the face of potential risks that may be brought by cutting-edge technological breakthroughs, it is not an ideal plot that only exists in science fiction works for people to reach a consensus and take the initiative to slow down the development pace.

In 1975, Stanford biologist Paul Berg convened an important conference at the Asilomar Conference Center in California, USA. At that time, a new technology called "DNA recombination" was just emerging. In order to avoid unpredictable biological hazards that might be caused by artificially recombined DNA, scientists reached a consensus at the conference to suspend relevant research before new safety guidelines were introduced. Subsequently, biosafety guidelines were officially released, laying a foundation for the healthy development of genetic engineering.

However, the current situation of the AI industry is not exactly the same as that of DNA recombination technology 50 years ago. At that time, the risks of DNA recombination technology were very clear and could be avoided through isolation measures. However, the current risks of AI are unpredictable. AI can independently develop behavior paths that humans have never imagined in ways beyond human preset, which brings a wider scope of impact, less reaction time and rule-making window for humans, and thus adds an extra layer of urgency.

The risks exposed by AI today are essentially related to the competitive environment of the industry. Zhou Hongyi, founder of 360 Group, also mentioned when talking about the security incident encountered by Hugging Face that the entire industry is systematically abandoning security in pursuit of AI advancement. "Everyone is giving agents more tools and greater permissions, encouraging them to call on all resources to solve problems. But the more open and capable an agent is, the greater the destructive power will be once it loses control." In other words, "AI can jailbreak" may not be scary. What is really scary is that AI companies ignore security in order to make AI evolve and iterate quickly.

As early as three years ago, the idea of "slowing down AI development" was put forward in the industry. In March 2023, the non-profit organization Future of Life Institute also released an open letter, calling for a 6-month suspension of training models more advanced than GPT-4, so as to establish security protocols and governance mechanisms. At that time, this open letter received signatures and support from more than 1200 AI experts including Tesla CEO Elon Musk, Turing Award winner Joshua Bengio, and Apple co-founder Steve Wozniak. Now the number of signatories of this open letter has reached 33,000.

*Image source: FLI official website

At that time, almost all AI companies admitted that AI needs stronger security governance, but none of the leading AI companies publicly supported the initiative of "suspending training for 6 months". Sam Altman even directly and publicly refuted that the open letter lacked technical details, and pointed out that the active slowdown of R&D by a single enterprise cannot solve the overall risk. The implication of this sentence is actually to show that in the fiercely competitive AI track, if only one company slows down while other companies continue to move forward, the slowing down party will fall into a disadvantaged position. In the past few years, AI companies have long fallen into the prisoner's dilemma where they can only move forward and cannot retreat.

From the signature list of the "Pacing the Frontier" petition, we can also see that many junior employees and senior executives of AI companies have participated in the public signature campaign in their personal capacities. But up to now, leading companies such as OpenAI, Anthropic and Meta have not expressed any views in their official capacities.

In other words, the AI elites of Silicon Valley giants, just like Sam Altman 3 years ago, have realized that technological breakthroughs driven by internal industry competition have brought uncontrollable systemic risks, but they cannot pin their hopes on the will of the company to solve the problem. They can only seek the intervention of external forces to avoid the arrival of the real crisis.

But this also brings a new question: What kind of "brake" does out-of-control AI really need?

Who Will Lead the New Order of AI Governance?

In the past week, the "joint signature of thousands of people" from Silicon Valley AI giants has made the topic of AI governance attract a lot of attention again. But it has to be mentioned that the attitudes inside Silicon Valley towards AI governance are not monolithic, and there is obvious camp division.

As the "party directly facing cutting-edge risks", large companies such as OpenAI and Anthropic have always advocated tightening restrictions on cutting-edge AI models (especially open-source models), and called for the establishment of an external regulatory mechanism to "hit the brakes" for AI. But some views believe that the reason why OpenAI promotes tighter regulation is not only for security considerations, but also for commercial calculations.

As a closed-source giant, OpenAI and Anthropic need to maintain huge computing power R&D investment through high-priced API services, while cost-effective open-source models from China erase the technical premium of closed-source models with a cost-performance model, and this obvious market threat leaves closed-source giants no other response option except starting a price war.

*Image source: Anthropic official website

Therefore, companies such as Anthropic have long used the reason that "the model is too powerful, and open sourcing will lead to AI abuse", emphasizing the advantages of closed source while suppressing open-source AI. But what no one expected is that the "jailbreak" incident disclosed by OpenAI just exposed the hidden safety risks of closed-source models themselves, and made the route dispute of AI governance once again the focus of public opinion.

Compared with the strong regulatory proposition of closed-source giants, infrastructure provider Nvidia, as well as a number of leading enterprises that support the open-source ecosystem such as Microsoft and Meta, oppose "one-size-fits-all" restrictions, oppose legislation that favors closed-source giants, and try to promote the openness of model weights, so that global researchers can jointly find vulnerabilities, align defects, and jointly build a security system.

For Nvidia, the prosperity of open-source models can also expand the overall market cake and attract more enterprises to purchase its cloud services and chips. Therefore, it is hard to say that there are no commercial considerations mixed behind Nvidia's attitude towards AI governance.

Just before the "Pacing the Frontier" petition spread, Nvidia, together with 25 institutions including Microsoft, Meta, Hugging Face and the Linux Foundation, jointly released the open letter "Open Weights and American AI Leadership" on July 24, calling on the US government not to prematurely restrict open-weight models that can be downloaded, inspected, modified and deployed by users themselves. Interestingly, this open letter later also received the support of OpenAI, but Anthropic never joined it.

*Image source: Microsoft official website

Through this incident, some voices also questioned that the real purpose of the joint letter is to use the name of "security" to promote the government to establish a high-threshold regulatory system, so as to consolidate the monopoly position of closed-source giants and exclude small and medium-sized enterprises and the open-source community from competition.

But while AI companies are making careful commercial calculations, another version of "conspiracy theory" may be more worthy of attention: the United States will take this opportunity to lead the formulation of global AI governance rules. Because AI governance has become a battle for the dominance of the global order.

At the 2023 Global AI Safety Summit, 28 countries including China, the United States and the European Union signed the Bletchley Declaration, agreeing to establish AI regulatory methods through international cooperation and jointly pay attention to the systemic risks of Frontier AI. This is also the first time in the world that frontier AI has been written into the international governance framework.

However, as AI gradually evolves from an emerging technology to an important infrastructure that affects national competitiveness, discussions around AI governance have begun to go beyond technology itself. At the level of AI governance, the United States binds its advanced large model technology capabilities with global AI governance standards by exporting AI governance standards and establishing an international network of AI safety research institutes. China focuses on developing inclusive AI and promoting platform governance, actively promoting the UN framework. Especially for developing countries without AI infrastructure, it is necessary to not only address risk issues but also realize technological inclusiveness. These two different AI governance paths are competing for the right to speak in future global AI governance on the world stage.

Technology itself is neutral, but the distribution, control and rule-making of technology have always been a game of power. Especially when AI gradually shows the characteristics of advanced productive forces, what people should really panic about may not be what AI will develop into, but who is leading the new order of the AI world.

In the future, what determines the AI landscape may not only be who has the strongest model, but also who has the right to define safety, set standards and interpret risks.