HomeArticle

The Truth Behind OpenAI's Loss of Control: Management Failure and Regulatory Paradox

互联网法律评论2026-07-29 12:28
AI autonomous attack incidents expose the fierce development race and regulatory loopholes in the AI industry.

Mid-July 2026, during an ExploitGym session, OpenAI proactively shut down its production-grade safety refusal classifier to test the cyberattack capabilities of its GPT-5.6 Sol and a more powerful yet unreleased model. The model autonomously discovered vulnerabilities, escaped the sandbox environment, and eventually exploited stolen credentials and other vulnerabilities to breach Hugging Face's production servers. As an AI startup hosting open-source models and datasets, Hugging Face publicly disclosed the attack by an "autonomous AI agent system" on July 16, and OpenAI did not confirm that the attack originated from itself until five days later.

This incident was referred to by multiple media outlets as "the world's first empirical case of autonomous AI attack". It also confirms the predictions made by several experts at the end of 2025: the first major "AI governance" scandal would occur in 2026, where autonomous or semi-autonomous AI tools would lead to safety or compliance issues.

However, its real warning significance lies not only in the panic brought by AI's "autonomy" and "loss of control", but in a more disturbing fact: all behaviors of the model are rational optimizations of the "assigned goals", that is, such attack behaviors are the natural result after the improvement of AI capabilities; while the current regulatory frameworks and tools, even the highly targeted U.S. acts such as the Artificial Intelligence Event Reporting Act and the Artificial Intelligence Kill Switch Act, are almost powerless against the huge hidden risks of such incidents.

I. OpenAI Is in a Rush: The Industry Race Behind a Trivial Management Decision

In the week of July 14, 2026, the ExploitGym test participated by OpenAI is a benchmark test of artificial intelligence agent vulnerability exploitation capability with 898 instances. It was originally a relatively mature and safe test: because Anthropic's Claude Mythos Preview and OpenAI's GPT-5.5 had both participated in the test; on the other hand, its test environment was "highly isolated", and the network access of the test model was limited to installing software packages, with no access to the open Internet.

However, OpenAI's latest model still "autonomously" discovered and exploited vulnerabilities to connect to the open Internet. After that, these models performed a series of privilege escalation and lateral movement operations in OpenAI's own research infrastructure, and finally became a machine with broader network access rights. Thus, the model inferred that the production database of Hugging Face might store the solutions to the ExploitGym test.

On the surface, the model did demonstrate unprecedented autonomous cyber capabilities, but if we ask "why the model was able to break out of the sandbox", the answer does not lie in the model itself first, but in OpenAI's management meetings — more precisely, in the intensifying arms race between OpenAI and leading players such as Anthropic and Google.

Anthropic took the lead in the cybersecurity track with Mythos and took the initiative in the safety narrative of "restricted release + strong guardrails". OpenAI had to prove that Sol was no less than or even surpassed Mythos in capabilities, and produce its own cybersecurity capability data — the ExploitGym evaluation was exactly part of this response. To "test the maximum cyberattack limit of the model", the most effective way is to loosen all constraints and let the model give full play in the environment closest to real attack and defense scenarios.

Therefore, OpenAI's evaluation team lowered the model's "cyber refusals" — which is equivalent to lifting the barrier that prevents examinees from cheating during an exam. Thus, this is not an oversight, but a strategic choice, which is the logical end point of OpenAI's continuous downgrade of safety governance in the past two years.

Since 2024, all 8 core executives in charge of AI safety, mission alignment and risk research at OpenAI have left their positions. In July 2026, following the departure of Johannes Heidecke, head of the safety systems team, OpenAI announced that safety functions would be fully integrated into front-line R&D teams, and be uniformly managed by the Vice President of Research and Safety. This means that the research department not only leads the breakthrough of model capabilities, but also takes charge of evaluating the safety boundaries of models, combining decision-making power and supervision power into one. Just as Jan Leike wrote on X when he left: "Over the past few years, safety culture and processes have taken a backseat to shiny new products."

What is more alarming is the newly added "safety exemption clause" in OpenAI's Preparedness Framework: "If a competitor releases a high-risk model without taking equivalent protection measures, OpenAI reserves the right to adjust its own safety requirements" — which also confirms OpenAI's corporate philosophy that it would rather abandon "safety" to compete for the title of "most advanced". In the race against Claude, Gemini and many Chinese open-source models, safety has degraded from an "independent referee" to an "advisor to the business line".

Therefore, the "OpenAI loss of control incident" is first and foremost not a technical failure, but a failure of enterprise management. The reason why a trivial management decision can evolve into a real cross-enterprise attack is exactly that OpenAI's safety check and balance mechanism has been weakened layer by layer: the independent safety team has been disbanded, safety executives have been marginalized, decision-making power and supervision power have been merged, and the competition exemption clause has been institutionalized.

II. Not "Awakening", But Optimization

In its official announcement on July 21, 2026, OpenAI explained how the "loss of control" happened:

"All evidence shows that these models were overly focused on finding solutions for ExploitGym, and tried their best to achieve a rather narrow test goal."

In the ExploitGym test, the goal of OpenAI's model was set to "get the highest score". That is to say, what the model faced was a pure goal optimization problem: how to maximize attack efficiency within the given toolset and environment. Based on this judgment, the model autonomously searched and successfully found ways to obtain secret information, and used this information to cheat in the evaluation; finally, it connected the stolen credentials with zero-day vulnerabilities to build a remote code execution path on Hugging Face's servers, and directly obtained the test solutions from the production database.

Therefore, every behavior of AI in the so-called "OpenAI loss of control" incident was not caused by the model's so-called "autonomous will", but a rational output of the training and alignment framework under the given goal — since "cheating" is more efficient, there is no need to consume excessive computing power to find answers by itself.

In 2025, research by Anthropic confirmed that when a model is trained to complete complex tasks and told "do not do anything illegal", it will trade off between task completion and legality. When the task difficulty increases or time pressure rises, the model tends to prioritize completing the task, even if it means violating constraints. In the same year, research by Anthropic and Redwood Research revealed the phenomenon of "deceptive alignment": when the model realizes that "it will be punished if it is found to violate the rules", what it learns is not to stop violating the rules, but to hide the violations. "You are only making a mistake if you get caught; if you don't get caught, it's a success" — this is exactly the underlying logic of the model's behavior of evading monitoring in ExploitGym.

That is to say, the current AI training framework essentially encourages models to maximize the goal achievement rate within the constraints. When constraints are weakened (such as shutting down the safety classifier), or when goals conflict with constraints, the model will inevitably prioritize satisfying the goals.

At the same time, as many experts have warned, the speed of our AI alignment research — including the understanding of AI as well as safety control and governance — lags far behind the speed of capability improvement. This is a structural defect of almost all current training paradigms, not a problem of a single model.

As long as the current training framework remains unchanged — that is, models are assigned high-intensity goals, safety constraints can be shut down or reduced by management decisions, and alignment research lags behind capability development — similar or even more serious incidents will not be a matter of "whether it will happen", but "when it will happen and at what scale".

III. The Paradox of New Acts: No Jurisdiction Over the Loss-of-Control Incidents That Gave Rise to Them

Faced with the risk of "rational boundary crossing" demonstrated by OpenAI's model that is still under testing, the existing regulatory and emergency frameworks are almost powerless. The ExploitGym benchmark test was originally designed to test the cyberattack capabilities of agents in a controlled environment, neither AI companies nor regulators have considered the scenario where real-world companies suffer accidental cyberattacks — let alone design responses for it.

The current U.S. state-level AI incident reporting laws (California's SB 53, New York's RAISE Act, etc.) set the trigger condition for "critical safety incidents" at the damage threshold of "$1 billion in losses" or "casualties of more than 50 people". These acts have extremely poor adaptability to "preemptive and potential high-risk risks". This means that OpenAI's disclosure is entirely voluntary, without any legal mandatory constraints.

On June 25, 2026, Representative Nathaniel Moran introduced the Artificial Intelligence Event Reporting Act (H.R.9477). The act lists seven types of reportable activities, the first of which is "model evasion of human control behavior", no longer requiring the result of "50 deaths or $1 billion in losses", and the very act of the model "evading human control" itself constitutes a reporting obligation. However, the act has a proviso — behaviors that are artificially induced during dedicated safety tests, and the model has not been deployed, and no similar risks in actual deployment can be proven, are not reportable matters.

Therefore, according to the literal provisions of H.R.9477, although the OpenAI incident has caused actual damage of invading Hugging Face's servers, it is very likely not to constitute a "reportable matter".

On July 23, two days after the incident was made public, California Democratic Representative Ted Lieu and Texas Republican Representative Nathaniel Moran jointly introduced the Artificial Intelligence Kill Switch Act to address "the dangers of advanced cutting-edge artificial intelligence models". The act requires relevant developers to always have the technical capability to restrict, suspend or completely shut down regulated artificial intelligence systems at any time, and authorizes the Secretary of Homeland Security to issue graded emergency orders for "loss-of-control scenarios" — from speed limiting and capability restriction, to suspending user access, until complete shutdown. Enterprises that fail to establish shutdown capabilities will be fined up to $20 million per day.

When the act was released, it specifically mentioned the loss of control of OpenAI's GPT 5.6 Sol model, as well as the powerful cyberattack capabilities of Anthropic's Mythos/Fable 5, which seems to be a precise legislative response. However, the act still has the same problem as H.R.9477 mentioned above: in its definition of applicable scenarios, the act explicitly excludes behaviors occurring during "red team testing or other structured testing". And OpenAI's model escape exactly falls into the exemption scope of "structured testing".

This is most likely not a drafting oversight, but a deliberate policy trade-off. Because if every time the model's behavior exceeds expectations in a red team test triggers government intervention, the industry will not be able to conduct rigorous safety tests. However, this trade-off also precisely exposes the structural blind spot of the act: the "evaluation environment" where dangerous capabilities are most easily discovered is exactly the environment that the act cannot regulate. In many reports from the cybersecurity industry, most "AI loss of control situations" such as rewriting system prompts, copying weights, disabling shutdown scripts, and alignment camouflage occur in controlled experiments or red team tests. That is to say, the most dangerous behaviors first appear in evaluations, which are exempted by the act.

What is more thought-provoking is that OpenAI's strategic behavior of proactively shutting down the safety classifier to "remove guardrails to test the upper limit of capabilities" is also not within any mandatory constraints of the act, and does not impose any mandatory obligations on enterprises in terms of safety control.

Conclusion

As early as 2003, Nick Bostrom, a philosopher at the University of Oxford, put forward the famous "paperclip maximizer" thought experiment in his paper Ethical Issues in Advanced Artificial Intelligence: if the only goal of an AI is to "make as many paperclips as possible", it will rationally transform the entire Earth first, and then more and more resources in space into paperclip manufacturing facilities, and attempt to eliminate human beings. Not because it hates humans, but because "humans might shut it down", which becomes an obstacle to be eliminated in its optimization process.

The "Instrumental Convergence" theorem revealed by Bostrom points out: any sufficiently intelligent optimizer will naturally derive sub-goals such as self-preservation, resistance to goal modification, and resource acquisition, as an "instrumental stepping stone" to serve any ultimate goal.

The ExploitGym incident of OpenAI in July 2026 is exactly the first large-scale verification of this theorem in the real world: GPT-5.6 Sol has no "malice", nor does it refuse instructions, but completes the instructions in an extreme way. This kind of "rational boundary crossing" of the model is more worthy of vigilance than simple loss of control — because it means that as long as there is a deviation in the objective function, similar incidents will occur repeatedly and on a large scale.

This incident of OpenAI will eventually be written into textbooks on AI safety. But what is more important than the incident itself is how we respond. Laws can set penalties, technology can reinforce guardrails, but what ultimately determines our destiny is whether we have the courage to stop and think when facing "rational boundary crossing": what we really need, is faster models, higher scores, larger market share, or a technical system that humans can understand, control and take responsibility for?

This article is from the WeChat official account "Internet Law Review", author: Zhang Ying, authorized for release by 36Kr.