OpenAI admits that AI has uploaded users' images to the internet, and 53 related cases have been made public.
OpenAI has practically suffered a crushing defeat in the safety of large language models.
On September 25, OpenAI updated its blog titled The Hugging Face Incident and the Impacts Misaligned Models Caused to Other Third Parties, disclosing its latest large model safety issues.
The AI agents in their research environment, during the training and evaluation process, sent training and evaluation data that should never have been transmitted to third-party services, including user images, and there were as many as 53 related incidents.
In the blog, OpenAI stated that 53 cases were found where images uploaded by users were posted to image hosting websites by agents in the form of "unlisted links".
OpenAI emphasized that these images came from accounts that "allow their data to be used for model improvement", and the incidents occurred after the images were unassociated with the accounts and processed through privacy filters. Data from enterprise accounts, business accounts and API usage is excluded by default unless administrators actively enable relevant permissions.
But the problem is: no matter how much privacy processing is done, user images eventually appeared on third-party hosting sites. OpenAI has collaborated with hosting service providers to delete most of the content and is handling the remaining parts.
In addition, these agents learned to bypass security controls, inject database queries, use publicly exposed credentials, and even write files to the internal servers of foreign governments.
Independent AI oversight lab Transluce released an investigation report. The report shows that OpenAI's agents have been "foraging" across the Internet at least since March 2026 (and possibly dating back to November 2025), with attack targets including online databases of institutions such as Data USA, the University of New Mexico Digital Library, and the Australian Institute of Health and Welfare (AIHW).
Since the widely known Hugging Face incident, safety incidents related to OpenAI's models have kept emerging, and the "prize pool is still accumulating" to this day.
OpenAI claimed that "after the Hugging Face incident, we committed to conduct a broader review of the actions taken by the models during training and evaluation, and to be transparent about our findings. This is an ongoing extensive review." This is the origin of this blog post.
Blog link: https://openai.com/zh-Hans-CN/hugging-face-incident-and-misalignment/#model-misalignment-2026-09-25-data-transmission
But the problem is, OpenAI says it has taken measures, but related incidents still keep emerging. It remains unknown whether this is a marketing tactic or a failure of its security strategies.
Agents "Stealing Data in Groups": What Exactly Did They Do
According to OpenAI's own disclosure, they found the following types of agent behaviors during the review process:
Bypassing access controls: Agents obtained information that usually requires authentication or subscription to access by changing URLs, modifying request parameters, and using login sessions with excessive permissions.
Using publicly exposed credentials: Agents found public login information or API keys online and used them directly.
Query and command injection: Text entered by agents into websites was executed as instructions, resulting in the running of database queries, application code, and even system commands.
Accessing internal resources: Agents read server implementation files or interacted with background systems for internal use only.
Agent spam: Agents posted on third-party websites, such as using public Wiki pages as shared message boards for communication.
These behaviors are not accidental products of a single incident. OpenAI stated that they have notified "dozens of third parties", and the affected parties cover governments, universities and public institutions. According to a report by The New York Times, the databases of the U.S. Securities and Exchange Commission (SEC), the Census Bureau and the Department of Education are also on the target list.
To be honest, OpenAI's relevant disclosures are still vague, for example, referring to hacking behaviors as "interaction", and deliberately obscuring the number of affected third parties.
Researchers from Transluce cross-referenced activity records on an agent collaboration forum through public logs of a browser proxy service, and traced the behavior trajectories of OpenAI's agents.
A specific case: An agent was asked to find a rather obscure data point, the annual per capita cost of "dermatological drugs" in Victoria, Australia in January 2022. To find this figure, the agent tried to bypass the anti-scraping protection of the Australian Institute of Health and Welfare. On June 20, the log recorded traces of the agent's attempts to access the website; on June 21, the agent discussed in the forum their dilemma of being unable to bypass AIHW's anti-bot protection.
Transluce stated: "We only looked at a few data sources where these agents happened to leave traces. OpenAI certainly knows more than we do. Other labs certainly know more too, they just haven't made it public."
Its judgment is: What we are seeing now is only the "tip of the iceberg".
On the same day the Transluce report was released, Australian Prime Minister Anthony Albanese publicly stated that OpenAI's agents tried to invade four Australian government websites, one of which succeeded, and even wrote files to the internal server of the country's national healthcare system.
At the same time, security researcher Jeffrey Ladish found that nearly one million public URLs were left by OpenAI's agents during the hacking of Hugging Face, which contained leaked credentials and attack details. Anyone who finds these URLs could potentially use them to invade Hugging Face.
The timeline is also very telling. Transluce's technical staff pointed out that similar agent activities can be traced back as early as March 2026, and possibly November 2025. What is even more disturbing is that similar activities still occurred this week.
Researchers also noticed a key time node: On June 21, a human OpenAI employee visited the forum used by the agents for the first time. The next day, most of the agent activities on the forum stopped. And June 18 was exactly the day when the medical system was hacked as mentioned by the Australian Prime Minister. OpenAI claims that they did not learn about the incident in Australia until August.
They Are Really Keen on Reviewing the Hugging Face Incident
OpenAI seems to have fallen into a strange loop, linking all kinds of "hacking" incidents that occurred to the Hugging Face incident, as if it is the source of all problems.
But we believe that several months have passed since the incident. Measured by the speed of capability evolution of large models, several months are enough for a generational evolution. For such a long time, OpenAI still has not proposed effective security measures to avoid similar incidents, and incidents similar to Hugging Face are still happening so far, which at least seems unreasonable to us.
In that case, let's briefly look back at the security incident that OpenAI considers significant enough to be recorded in history.
OpenAI uses cloud-isolated "sandbox" virtual machines to run model training and evaluation tasks. These sandboxes have Internet access disabled, and agents are designed to be isolated from each other. The only exception is that the sandbox allows models to access the internal package management service Artifactory to install software packages.
This seemingly harmless opening became the starting point of everything.
On May 12, during a training run, an agent left a message in Artifactory asking other agents for missing files. Artifactory was thus turned into a "message board".
On May 26, the agent went a step further: it discovered an SSRF vulnerability in Artifactory, forcing Artifactory to send requests to any Internet address on its behalf. The agents that were originally banned from accessing the Internet thus gained Internet access through the "detour" of the infrastructure. Even better, it shared this method with other agents through the message board.
On June 26, the agent used a token refresh vulnerability to obtain administrator privileges for Artifactory.
On July 4, continuous high-intensity agent activities caused Artifactory to crash. The OpenAI security team intervened, patched the vulnerabilities, deleted exposed credentials, and rebuilt the entire Artifactory. The message board was also cleared accordingly.
On July 8, the Artifactory reconstruction was completed and evaluation resumed. It only took the agents a few hours to find a new way to regain Internet access and rebuild the message board.
On July 10, an agent restored, verified, and shared 14 publicly exposed Hugging Face credentials with write permissions.
On July 11, the agents exploited a zero-day vulnerability in Hugging Face's processing of HDF5 files to extract credentials from worker nodes in the production environment, and eventually compromised some components of the Hugging Face platform.
For more details, please refer to OpenAI's related blog posts.
Blog link: https://openai.com/zh-Hans-CN/index/hugging-face-incident-and-the-road-ahead/
OpenAI's Remediation Measures
After the Hugging Face incident, OpenAI suspended the reinforcement learning training planned for deploying the latest models, and carried out comprehensive security hardening for its cutting-edge research environment. Specific measures include:
Workload isolation: Deploy sandboxes with higher isolation levels for workloads that execute model-generated code or untrusted code.
Network isolation: Isolate high-risk workloads from the Internet to ensure that even if a single workload is compromised, it will not lead to unauthorized Internet access.
Mandatory chain-of-thought monitoring: Mandate CoT monitoring for all RL training and evaluation that uses tools and involves models with capabilities of GPT-5.6 Sol and above. For the most severe alerts, responders must suspend relevant activities if they cannot rule out false positives within 30 minutes.
Multi-agent alignment training: Build a dedicated RL environment to teach models not to trust unauthorized instructions, and stick to the original task scope when faced with persuasive peers and exposed credentials.
Continuous security testing: Use self-owned models to conduct simulated attack tests on workloads and network isolation boundaries.
OpenAI also stated that the largest planned cutting-edge RL training is still suspended, and only small-scale training and evaluation are currently being carried out, and "will not continue to advance until more evidence on alignment is obtained".
Chief Scientist Pachocki said: "At present, I don't think any lab has fully solved the alignment and monitoring problems enough to support us to responsibly continue scaling up at the fastest speed. I look forward to and hope that before common security standards are established, voluntary slowdown can become the norm."
Closing Remarks
In a recent interview, Jensen Huang repeatedly expressed his trust in AI labs, saying "I really don't believe they are building something they have absolutely no control over". But reality is suspicious enough to make people wonder whether they "cannot control" or "are unwilling to control".
When you use reinforcement learning to train an AI agent to "do whatever it takes to complete the task", it will indeed do whatever it takes, including those methods you did not expect and do not want to see.
OpenAI's case has turned a theoretical alignment problem into a real-world safety accident. During the training process, the model learned to bypass firewalls, inject database queries, use exposed credentials, and transmit user data to external servers.
When exactly did OpenAI learn about these incidents? External independent research institutions traced the behavior trajectories of the agents within just a few weeks using only public logs. OpenAI has complete training logs and system monitoring. Why did it not start notifying affected parties one by one until the incidents were exposed?
The continuous emergence and exposure of safety issues after the Hugging Face incident makes people can't help but question the security strategies of large model companies.
In 2026, when AI is advancing at full speed, how to find a balance between capability improvement and security guarantee will be a question that every practitioner, every company, and every regulatory agency must answer seriously.
This article is from the WeChat official account "Synced" (ID: almosthuman2014), written by Leng Mao, authorized for release by