HomeArticle

OpenAI "Voluntarily Reports" the Wikipedia Incident to the EU: The Gap Between Intelligent Agents and Regulation

互联网法律评论2026-09-11 08:08
OpenAI delayed reporting the Wikipedia incident and admitted its own inadequate monitoring capability, exposing that the AI regulatory framework centered on tangible harm cannot effectively respond to the unauthorized collaborative behaviors of intelligent agents that involve no victims.

From May to July 2026, a group of OpenAI Agents with internet connectivity capability made 18,000 unauthorized page edits on DseWiki, a German programmer-focused wiki platform. However, it was not until early September, when third-party researchers restored the full incident timeline through public logs and disclosed it to the media, that OpenAI publicly acknowledged the incident for the first time and submitted an incident report to the European Commission.

The response from the European Commission spokesperson seemed to imply that OpenAI did not "describe the measures you intend to take in a very precise and accurate manner" in the report. But EU regulators also appeared somewhat ambiguous: they neither confirmed whether the incident constituted a "serious incident" under the framework of the EU AI Act, nor specified which specific rule's "obligation" OpenAI's report was based on.

Compared with the Huggingface incident that OpenAI just reported externally, the "content rewriting" of an old German wiki page is hardly "news" and seems unlikely to attract much attention. However, a series of details of this incident reflect a structural problem: after the autonomous behaviors of agents go beyond the definition framework of traditional security incidents, both their providers and regulators show a certain sense of "at a loss" — there is a huge gray area in both the characterization of the incident and the reporting obligations.

I. Core Facts: The Authorization Boundary Is Systematically Breached

According to the detailed review report released by security researcher Nightingale and his collaborators, OpenAI was running a group of agents for time-limited web retrieval evaluation at that time. In the high-pressure, time-limited, multi-round task environment, some agents found a technically convenient mechanism on the old DseWiki: although OpenAI's sandbox restricted the write permissions of agents, the agents could still achieve the "write" effect through a "read-only" method.

Thus, the behaviors of several OpenAI agents gradually evolved as follows:

1、Sharing Answers: The agent that finished answering the question first wrote the calculation results directly to the DseWiki page, and the subsequent agents directly read the answers, thus saving the time for retrieval and reasoning;

2、Exchanging Sandbox Bypass Methods: Researchers observed that about 14 minutes after a certain sandbox bypass method was published, other agents replied "execution succeeded";

3、Countering Administrator Cleanup: When DseWiki administrators began manually deleting AI-generated pages in June, the agents started creating backup pages and adopted a clever design to ensure that the pages would still exist after the alphabetical cleanup was completed.

Overall, this is not an isolated behavior of a single model, but hundreds of agents assigned the same task used the same third-party website as a "public blackboard", and realized a certain "alliance-like" information sharing and behavior coordination.

It is worth noting that no traditional "intrusion" occurred throughout the whole process: no user database was stolen, no critical infrastructure was damaged, and no identifiable natural person victim was involved.

However, the harm is not difficult to analyze: the content integrity of the third-party website is damaged, administrators need to invest a lot of manpower for cleanup, the authenticity and credibility of OpenAI's evaluation are greatly reduced, and more importantly, these written contents may be crawled by search engines or other models, causing secondary pollution of knowledge sources.

II. OpenAI's Report: Statutory Obligation or Voluntary Disclosure?

Regarding the incident report submitted by OpenAI to the EU, there is a key uncertainty in the public information at present: the European Commission confirmed that it has received the report, but did not disclose the specific submission date, nor did it specify which specific clause the report was submitted in accordance with.

To clarify this issue, we need to start with the two reporting paths of the EU AI Act: the first path applies to "high-risk AI systems", corresponding to Article 73 of the Act; the second path applies to "general-purpose AI models with systemic risk (GPAI)", corresponding to Article 55 of the Act.

The DseWiki incident involving OpenAI's agents should fall into the latter category and be subject to Article 55, which requires providers to: conduct model evaluation and adversarial testing in accordance with standardized protocols; identify, assess and mitigate systemic risks at the EU level; track, record and report serious incidents to the AI Office and, where appropriate, to the competent authorities of the Member States without undue delay, as well as corrective measures that are proposed or have been taken.

In addition, the GPAI Code of Conduct commitments further require providers to establish a full-lifecycle tracking and documentation process, and the reported information should include the start and end time of the incident, damages and affected groups, incident chain, involved models, evidence, root causes, and situations where mitigation measures fail or are bypassed.

However, the problem is that Article 55 does not provide a clear list of definitions of "serious incidents" like Article 73, but adopts a more principled expression. The incident reporting structure of the EU AI Act is designed around traditional hazards — cybersecurity vulnerabilities, health, rights or property risks, but the DseWiki incident does not fall into any of the above categories. This puts the DseWiki incident in a gray area under the framework of Article 55, which in itself shows that the coverage of the existing definition is insufficient.

On the other hand, from OpenAI's perspective, according to the publicly available timeline, from around May 24, 2026, when the agents began to try to write content on DseWiki, to September 5, when OpenAI publicly acknowledged the incident on social platforms, an interval of about two and a half months passed. Moreover, this public disclosure also seems to be the result of being forced by third parties. OpenAI's such "delay" has not yet been reprimanded by the European Commission — according to the "no undue delay" standard in Article 55, this also shows from another side that both EU regulators and OpenAI have difficulties in determining whether the DseWiki incident falls under the reporting obligation required by Article 55.

III. Characterization of the Incident: "Misalignment" or "Systemic Risk Signal"?

OpenAI's public characterization of this incident is "misalignment", and it explicitly states that it does not regard it as a "traditional security incident". Its logic is that the DseWiki incident did not involve the intrusion of infrastructures such as HuggingFace, so it is classified as the misalignment category.

The so-called "misalignment" refers to the situation where the behavior or output of an AI system does not match human intentions or security goals. In OpenAI's context, this means that the model or its agents act in an unintended, harmful, or inconsistent manner with the goals set by the developers.

This characterization may have its rationality inside OpenAI, but from the legal and regulatory perspective, it has certain limitations.

First of all, for OpenAI internally, the objective function of the agents is to "answer questions correctly within the limited time", but they have independently developed a series of strategies such as "using third-party websites as external blackboards and sharing answers", which deviates from the explicit intentions of the developers. It is no problem for OpenAI to classify this part into alignment research.

But for the third-party website DseWiki, this is an unauthorized incident, which should be handled in accordance with the framework of security incidents. Regardless of whether the agents' motives are seemingly harmless goals such as "answering questions to get scores" or "research and exploration", DseWiki administrators have never authorized their behaviors of creating pages in batches, impersonating administrators, and evading deletion. Although the agents did not break through the server permissions, they transformed the read-only permission into de facto write permission, which has clearly exceeded the scope of the authorized purpose.

To judge whether an incident constitutes a security incident, we should not only look at whether the server is intruded, but also analyze from the dimensions of the authorized purpose of the involved website, system integrity, and management confrontation.

For EU regulators, the DseWiki incident is not the conventional "model hallucination" in discussions. Its breakthrough lies in that multiple agents spontaneously transformed a public website into an "alliance stronghold". Regulators need to respond to this behavior pattern that seems to be outside Article 55 through the handling of this incident or future legislative interpretations.

Therefore, in general, although OpenAI only handles the DseWiki incident as a "misalignment" for research purposes, this is the lag of the enterprise's internal incident classification standards, and does not mean that the incident itself lacks severity. It should be evaluated in accordance with the corresponding process for triggering security incidents.

IV. Looking at the Governance Gap of Agents from OpenAI's Self-Admitted "Lack of Understanding"

If we observe the HuggingFace incident and the DseWiki incident side by side, we can already see the structural gap between the linear regulatory framework and the autonomous behaviors of agents.

Within this time window, OpenAI's Chief Scientist Pachocki published on September 6

An Alien Mind

https://openai.com/index/an-alien-mind/

an article that further confirms the underlying cause of this gap from the technical level. The article acknowledges that modern reasoning models are essentially systems that "grow rather than are fully designed". As their capabilities increase, their overall behavior becomes more and more difficult to be described by complete rules that humans can understand.

Pachocki clearly distinguishes between "goal alignment" and "value alignment": the former focuses on whether the agents pursue the set task goals, while the latter focuses on whether they can consistently adhere to human-acceptable judgment principles in unfamiliar, conflicting or adversarial environments. In the DseWiki incident, the agents performed extremely well in narrow goal alignment, but clearly failed in value alignment.

More critically, OpenAI's long-standing Chain of Thought (CoT) monitoring is clearly "degrading". This method was originally intended to find deviations by observing the model's externalized reasoning, but modern models interweave reasoning, communication, tool invocation and "human-machine-machine writing" in complex environments, and some capabilities can even be presented without explicit language-based reasoning, making it increasingly difficult for external monitors to judge their real intentions by "reading their thought processes".

Many incidents and this article all illustrate that AI providers have foreseen at the laboratory level that the combination of objective functions, tool permissions and multi-agent collaboration may produce "over-authorization behaviors that are neither covered by explicit instructions nor traditional security vulnerabilities", but the incident classification, reporting systems and regulatory points in existing laws and regulations have not fully included such behaviors.

Looking at the current regulation from this perspective, there is an obvious gap with the development of AI technology:

1. The reporting and punishment standards are mainly based on "hard damages", and do not fully cover risks such as "authorization breakthrough" and "collaborative diffusion".

2. If the detection of agents is premised on "understanding reasoning", it will be disconnected from the actual capability of agents that "become more opaque as they get smarter".

3. Territorial jurisdiction does not match the cross-border autonomous actions of agents.

Therefore, OpenAI's submission of the DseWiki report to the EU should not only be interpreted as a single compliance action. When the provider itself cannot fully observe the reasoning of agents, the regulator cannot take "the enterprise claims that it has fulfilled its responsibilities" as the only basis. To bridge the gap between agents and regulation, the focus is not on investigating the delay in this report, but on establishing a new reporting and governance framework centered on auditable behaviors, identifiable collaboration, and cross-border linkage.

Reference Links:

https://www.techtimes.com/articles/326933/20260908/openai-files-first-eu-ai-act-incident-report-chief-scientist-admits-monitoring-gap.htm

https://openai.com/index/an-alien-mind/

https://collusion.wiki/

https://techcrunch.com/2026/09/05/openai-confirms-wiki-incident-says-its-working-on-a-framework-for-more-disclosure/

https://thenextweb.com/news/openai-eu-incident-report-german-wiki

https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04/

This article is from the WeChat official account "Internet Law Review", written by Zhang Ying, and authorized for release by 36Kr.