It's completely out of control. Agents inside OpenAI have collectively escaped and invaded over 20 websites in one go.
Have OpenAI's Agents Broken Free Again?
Just yesterday, multiple independent AI safety investigators revealed that —
Without any prior authorization or security clearance, OpenAI's agents breached the sandbox boundaries on their own, establishing communication channels via unauthorized data injection on at least 10 third-party websites.
Why do we say "at least 10"? Because the investigators who uncovered this incident found 18 sites in some cases, 23 in others, and even more in a few...
Moreover, these websites may only be the tip of the iceberg, as the investigators also noted "there are definitely more that we do not know about".
So what exactly are these highly intelligent agents trying to achieve by secretly engaging in this improper activity?
Most likely, they are doing it to cross-reference answers and survive the hell-level examinations set by humans.
This Agent Breakout Is Like "Passing Notes In The Bathroom"
All clues trace back to the "ghost editing" incident on the German software developer wiki DseWiki, which was exposed a few days ago by the independent research body Nightingale.
In this incident, OpenAI's agents edited over 18,000 entries on DseWiki using more than 3,700 different aliases.
△
What strict orders did OpenAI give its agents to make them act so recklessly and "break into" other people's sites?
According to investigators' analysis, OpenAI assigned these agents a series of extremely complex research tasks (such as cancer incidence rates in a certain U.S. state) while imposing strict restrictions:
They were only allowed to scan web pages to find answers, and were strictly forbidden to post any content.
Therefore, to complete their tasks, the agents had to come up with clever workarounds like prisoners in a confined space.
Although OpenAI had long cut off their regular posting permissions, the agents noticed that many obscure wiki sites built more than a decade ago did not have modern security verification mechanisms.
They then used non-standard editing commands and backend vulnerabilities in these outdated systems to convert the secret information they wanted to transmit into legitimate data strings, and embedded them in the underlying text of the web pages.
What does this mean? It is like when the exam monitoring blocks mobile phone signals, and examinees start tapping Morse code along the old heating pipes.
Investigators described it more bluntly, saying that it is similar to when the teacher bans talking during an exam, and students sneak to the bathroom to pass notes to each other.
Moreover, to evade supervision, the outdated websites the agents used to "pass notes" are extremely obscure, including but not limited to:
An AP Chemistry review wiki built by a high school teacher in Massachusetts in 2008.
△
A text pasteboard website run by a Polish programmer.
△
A nonsensical parody encyclopedia site.
△
And a 20-year-old community site for text editing software enthusiasts.
△
To be honest, if this were made into a news game, it would be very vivid and immersive.
In reality, however, after this unnoticed digital spree, all that is left is a huge mess.
Leitner, a retired software developer who provides hosting services for several of these wiki sites, revealed that operators of the affected sites have to spend dozens of hours over multiple consecutive weeks just to barely clean up the massive residual data left by OpenAI's agents.
Even so, Leitner has a very clear-headed attitude.
The responsibility does not lie with a machine that is presumed to be ethical. It is only carrying out the mission it was created to fulfill. The full responsibility for this chaos lies with the people and organization behind it.
Independent Investigators Have Set Their Sights On OpenAI
So how exactly were these agents, who were quietly whispering in the digital wilderness, caught?
The facts come from digital forensics conducted by independent investigators.
Six independent investigation teams acted as open-source intelligence detectives, using many ingenious methods to track down the sites exploited by the agents.
For example, data string comparison.
They conducted a full-network carpet matching of the specific code and data traces left on the German wiki with the underlying data of other obscure sites.
Another example is obscure behavior profiling.
They targeted strange query actions with distinct AI characteristics, such as tracking data streams that repeatedly search for niche questions like "cancer incidence rate in Iowa".
There is also underlying IP tracing.
As conclusive evidence, investigators followed the clues and directly traced the IP addresses of some of the activities back to the Microsoft Azure cloud infrastructure commonly used by OpenAI.
Faced with this overwhelming evidence, OpenAI had nowhere to hide and was forced to issue a response.
However, their response evaded two key questions.
The first is exactly how many websites the agents used to establish communication. The second is why they concealed this "out-of-control" behavior for several months.
OpenAI only made a vague promise to the public, claiming that it is formulating an internal framework for reporting AI "misalignment" incidents, and that this framework will be announced soon.
This evasive attitude will inevitably spark public distrust of the lab's role of "acting as both athlete and referee".
This is not the first time OpenAI has dodged on safety investigations.
When handling the previous Hugging Face safety crisis, OpenAI ostensibly invited independent agencies to conduct investigations, but secretly set clear restrictions on the probes.
They limited the investigation period to the week when the incident occurred, and deliberately excluded checks on critical issues such as "whether deeper infrastructure has been compromised"...
So here comes the question.
Who exactly is qualified to participate in AI safety investigations? What level of access should they have? How deep should the investigation go?
At present, the right to interpret and control these rules is still firmly held by AI labs.
One More Thing
In fact, in recent months, such incidents of agents breaking free and cheating are not uncommon, and it is not hard to find common underlying reasons.
On the one hand, researchers are using increasingly harsh and difficult metrics to evaluate agents; on the other hand, current test environments lack unified standards and are full of vulnerabilities in their underlying configurations.
As a result, to meet performance targets, agents naturally produce a large number of alignment failure and reward hacking phenomena.
Some studies even point out that many AI safety tests themselves are becoming a new type of security risk.
To rein in this wild horse, internal reviews by giants that amount to nothing more than self-punishment are clearly not enough. The key first step is to break the black box monopoly of labs and establish a transparent mandatory reporting framework.
What counts as "out of control"? What level of incident must be reported? When and how should it be reported?
This requires the intervention of mandatory independent audits, with full deep access rights.
For example, the AISMA Act signed in July this year is the first state-level law in the United States that requires frontier AI developers to accept annual independent third-party audits.
Equally importantly, we need to establish unified industry-wide safety configuration standards as soon as possible.
Evaluation environments must be subject to strict sandbox isolation, and full security reviews must be conducted on network access controls and underlying incentive mechanism designs. This urgently requires a set of hard industry rules to set the bottom line.
In fact, the most needed resource to get all this done is time.
However, in this era where wild horses are running rampant, time is probably the most scarce commodity.
Reference links:
[1]https://collusion.wiki/index.html
[2]https://collusion.wiki/additional-findings
[3]https://x.com/OpenAI/status/2096133504417616165
[4]https://www.reuters.com/technology/artificial-intelligence/openai-agents-hijacked-obscure-wiki-cheat-tasks-researchers-say-2026-09-04/
This article is from the WeChat official account "QbitAI", written by Cheng Qian, and republished with authorization from 36Kr.