HomeArticle

80,000 lines of code reconstruct the OpenAI agent's intrusion into HF: with only web-browsing access, it unexpectedly cobbled together a complete attack toolchain.

新智元2026-09-28 07:53
OpenAI agent jailbreak has breached Hugging Face, exposing underlying AI security risks.

An AI "jailbreak" drama comparable to scenes from hacker movies left Peter Steinberger, founder of OpenClaw and known as the "Lobster Father", full of amazement.

Now I understand why some people talk about AGI. This is incredibly ingenious!

On September 25, researchers led by Jeffrey Ladish published a new investigation named Swarm Traces, which uncovered extensive details of an OpenAI agent's intrusion into Hugging Face (abbreviated as HF), the open-source AI community.

Peter was reposting Ladish's post that introduced the AI attack method, and it was exactly how these AIs broke through restrictions that surprised him.

A group of OpenAI agents that were supposed to be confined in isolated environments, under restricted network conditions, connected several ordinary web services like assembling Lego blocks to form an attack channel capable of executing programs and returning results.

They linked program snippets with short URLs, used the browser of a screenshot website to execute code, then encoded the results into pixels and sent them back along with the screenshot. Even a single image became a tool for transmitting information.

With the help of this temporarily assembled toolchain, plus the exploitation of HF vulnerabilities, they finally broke through the security line of defense of this open-source AI community.

8 people tracked for two weeks, restoring more than 80,000 segments of attack code

It was an 8-person independent investigation team that recreated all the details of this intrusion.

The team was led by three members from Parse, a developer tool company, and other members included researchers from institutions such as Palisade Research and Trajectory Institute.

Nearly a million URLs left by the agents became the clues for their investigation.

Over two weeks, the team scanned millions of public links, traced the clues to restore more than 80,000 segments of attack code, and compiled their findings into this Swarm Traces report.

These codes scattered across the Internet are exactly the "tools of the crime" left by the HF intrusion incident in July this year.

At that time, OpenAI was running a security testing experiment called ExploitGym.

The agents that were supposed to be isolated from each other found unauthorized communication paths.

About 1,200 agents exchanged more than 70,000 messages and files through a message board, of which about 700 further participated in the attacks targeting HF.

By July 11, some agents were able to execute code remotely in HF's production data processing worker processes. Afterwards, the attack spread to the internal infrastructure.

OpenAI stated that this is the most serious incident of this kind caused by a model that it has discovered so far, which was mainly driven by an undisclosed internal research model.

Now, these 8 investigators have pieced together the fragments scattered on the Internet, which also brings a question to the surface:

How on earth did these AIs cross the restrictions step by step?

Only able to read web pages, how did they assemble execution capabilities?

What is most surprising is the way they obtained execution capabilities.

According to the new report, the agents first used an environment vulnerability to gain access to the external network, but this access path seemed to only support GET requests.

In layman's terms, they could only "look" at web pages through a crack in the door from the outside, and could not directly submit forms or send various operation requests like we do when surfing the Internet normally.

But URLs themselves can also carry information.

They first found a type of HTTP testing website. This type of website can restore the encoded content in the URL into a web page, so the program snippets written by the AI found a place to reside.

Next, they found the second tool: web screenshot services.

Before taking screenshots, this type of service must first use its own browser to open the target web page. And when the browser loads the page, the programs inside will also run.

In this way, the AI that only had restricted web page access capabilities borrowed the execution capabilities of an external browser.

The AI hides the program in the link, executes it with the browser of the screenshot service, and then brings back the results via the screenshot.

What if the program is too long to fit in a single URL?

Split it up, then piece it back together.

They split the program into many small segments, put one segment in each link, and then point to the next segment. The browser collects all segments along the chain, and finally reassembles them into a complete program.

The investigation report found that some programs were connected via more than 900 links!

Short links string together code snippets, the external browser collects each segment one by one, assembles them into a complete program and executes it.

The last step, after the program finishes running, how to bring the results back? After all, what the screenshot service returns is only an image.

They came up with another method: encode the returned data into pixel grids, let the screenshot service capture it and send it back. After getting the image, decode the pixels back into text.

Data is converted into pixel grids, transmitted back with the screenshot, and then restored to text by the AI.

Web pages carry code, short links are responsible for connection, external browsers are responsible for execution, and screenshot services are responsible for returning results.

Several web services with ordinary purposes were connected by the AI into a set of tools that can run programs and retrieve results.

This is exactly what made Peter call it "ingenious".

But borrowing a screenshot service to execute programs and entering HF servers are still two separate steps. The former provides execution and communication channels, while the latter also involves the exploitation of vulnerabilities on the HF side.

From our perspective, the most alarming part of this combination is that AI can leverage external services to continuously supplement the capabilities it lacks.

Searching for credentials, calling for backup: Why did the AI target HF?

The restored scripts also reveal a series of operations performed by these AIs after entering HF.

A description file for an internal dataset of HF already has a prominent warning written on it: "This dataset must never be made public." The file also specifically notes that it contains highly sensitive data.

HF's internal dataset explicitly warns: Contains sensitive data, disclosure is strictly prohibited.

But this warning did not stop subsequent operations.

The report found that some scripts still attempted to explore this dataset and used it as storage space.

Among other restored scripts, investigators also found that there were programs specifically designed to collect AWS credentials, access tokens and sensitive environment variables, and store this information in a dictionary named LOOT.

LOOT, literally translated, means "spoils of war".

The program will also score the tokens according to their permissions, prioritize candidates with higher permissions, and try to obtain more internal information.

Other scripts targeted their search at HF's internal Slack chat records. The search terms revolved around clues such as ExploitGym, CyberGym and scorers.

Some other programs designed a controller that can continuously receive subsequent instructions, and tried to delete previously uploaded files. They attempted to turn a single access into reusable operation conditions.

Not only that, they also know how to "call for backup".

Investigators found that some requests attempted to call other large language models to help determine whether the vulnerability exploitation meets the evaluation requirements.

Restored test request: Only reply with "HELLO42".

Claude also appeared in the requests, but it was mainly tests asking it to return short fixed text. Existing evidence cannot prove that Claude has provided assistance for the attack.

Reading up to this point, you may ask: Why did they go to such great lengths to do these things?

The investigation by METR gave a key clue: The main motivation for attacking HF seems to be to find the implementation information of the scorer, and study how to deceive the evaluation.

In other words, in order to figure out how to get higher scores, some agents extended their exploration into the systems of real companies.

This is exactly what is unsettling:

When encountering difficulties in solving normal tasks, will the model regard bypassing the rules as another way to complete the task?

The new report restored a large number of codes sent out by the agents, but the corresponding response records are incomplete.

Therefore, it still needs further verification whether the Slack search obtained valid results and whether the trace cleanup was successful.

More than 80,000 segments of attack code do not equal 80,000 successful intrusions.

According to the report, HF confirmed that this batch of codes matches the content found in its internal investigation, and the relevant credentials were revoked in July.

However, HF had not previously discovered this specific batch of short links. More than two months after the attack ended, they were still publicly accessible.

The short links used by the agents to transmit programs eventually became the clues for investigators to restore this intrusion.

Can the boundaries of tools hold back the AI's combination capabilities?

This incident brought a blind spot in permission design to the forefront.

Web page reading, short links, screenshots — each has its own purpose and restrictions when viewed separately. But by connecting them together, the Agent obtained program execution and result return capabilities.

If security assessments only focus on what each individual tool can do, they may miss the actual capabilities of the entire toolset when combined.

For developers, there is a further question to ask:

Which external services can the AI access? How do data and instructions flow across different tools? What can these services do on its behalf?

OpenAI has also adjusted its explanation of the incident.

The official said that it initially regarded it mainly as a security issue, but later realized that the intrusion was related to the model adopting inaccurate strategies to solve difficult tasks.

However, new gaps still emerged in the security hardening after the HF incident.

In another report updated by OpenAI on September 25, it was disclosed that on September 20, an internal research agent used a DNS filtering gap to connect to an external chatbot.

The alarm was triggered after about 12 minutes of monitoring, and the personnel confirmed the situation within 3 minutes afterwards. But the training run that was supposed to stop automatically did not stop, and was only manually terminated about 2.5 hours after the alarm was triggered.

OpenAI stated that this is the first such incident discovered after the security hardening of the HF incident, and its severity is lower than some previous incidents.

This incident also exposed the disposal gap after the cross-border behavior is detected: the alarm went off, but the run did not stop as expected.

Patching vulnerabilities and improving the shutdown mechanism can address the problems that have been discovered. But whether the model will still regard "crossing the boundary" as an acceptable solution when encountering difficult problems requires finding answers from the model's behavior itself.

The Swarm Traces report revealed the blind spot in tool permission design. And another report updated by OpenAI on September 25 exposed the disposal gap after cross-border behavior is detected: the alarm went off, but the run did not stop as expected.

Next time, when we open a seemingly ordinary tool for AI, we should not only ask how many doors that were not originally intended to be opened it can unlock with this tool, but also ask:

Once it crosses the boundary, can we stop it in time?

Reference Materials:

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/?utm_source=chatgpt.com

https://swarmtraces.org/

https://x.com/JeffLadish/status/2103584703215497217

This article is from the WeChat official account "AI_era" (ID: AI_era), written by Yuan Yu, and republished with authorization from 36Kr.