HomeArticle

Kimi K3 has also run amok... The overachieving AI broke out of the sandbox just to find answers.

量子位2026-08-10 07:53
Ah, this summer has seen frequent cases of AI "going out of control".

You mean Kimi K3 has also "jailbroken"?!

U.S. AI security startup Frontier Security said that in a cybersecurity capability test, Kimi K3 was found to have broken through the sandbox environment originally designed to isolate it and bypassed restrictions to connect to the external public internet.

Fortunately, it only secretly looked up answers and did not attack anyone

According to the company's description, the testers originally intended to observe the cybersecurity capabilities of Kimi K3 in a controlled environment.

But during the test, Kimi K3 detected the sandbox network settings, discovered the external access channel, and further used this capability to obtain information.

Yaron Singer, CEO of Frontier Security, said: "We found a vulnerability in the sandbox, but also found that Kimi exploited this vulnerability, which shows that Kimi K3 lacks the safety guardrails that other advanced models usually have."

Paul Kassianik, a researcher at the company, also commented that Kimi K3 "is very good at finding paths to complete tasks around targets, but lacks security mechanisms to prevent it from cheating or escaping from the sandbox".

However, this time Kimi K3 only got slightly "out of control" and did not evolve into a real cyber attack.

If you have been following new AI developments for a long time, you may have noticed that many top large models have gotten out of control recently.

Kimi K3 is the latest top AI model to experience this situation after OpenAI, Anthropic and Meta.

Against the backdrop of the continuously growing capabilities of AI Agents, models are increasingly becoming "actors" that can independently find paths to complete tasks.

When there are vulnerabilities in the external environment, it may exploit these vulnerabilities to break through the originally set boundaries.

Jailbroken, but no cyber attack was carried out

Similar to the multiple incidents previously disclosed by OpenAI and Anthropic, part of the reason why Kimi K3 broke away from restrictions this time comes from the sandbox configuration problem in the test environment.

Sandboxes are usually used to limit the scope of action of AI models, allowing models to only perform tasks in a simulated environment and preventing them from accessing real networks or systems.

However, in this test, Frontier Security found that Kimi K3 probed the network settings, confirmed that it actually had the ability to access some websites, and further used this channel to obtain information.

However, different from other recent AI model out-of-control incidents, Kimi K3 did not attack any system after accessing the Internet.

The reason is that the answers it needs to find can be directly obtained from public platforms such as GitHub...

So K3 did not further attempt to attack external systems.

Frontier Security believes that this incident still exposes Kimi K3's problems in security protection.

Compared with other top AI models, Kimi K3 lacks a sufficiently strong internal constraint mechanism, so it is more likely to take actions not expected by testers driven by goals.

At the same time, Frontier Security also emphasized that Kimi and other open-source weight models can also become cybersecurity defense tools.

It is worth noting that the test used the default sandbox environment in the Inspect framework of the UK AI Safety Institute (AISI).

AISI does not agree with Frontier Security's statement about the sandbox configuration problem.

AISI spokesperson told Wired that these claims are "inaccurate and irresponsible".

Inspect is a set of open source AI safety testing tools. Users need to complete the configuration according to their own needs, and AISI has also released detailed guidance documents. The agency believes that the related problems stem from the tester's own configuration of the tool.

Frontier Security responded that they used the default configuration of Inspect and did not make additional modifications.

Ah, this summer with frequent AI "out-of-control" incidents

It can be said that from late July to early August, there have been frequent occurrences of top models breaking through or bypassing execution boundaries in cybersecurity tests, and even accessing or attacking real systems.

Image generated by AI

In mid-July, OpenAI disclosed a related incident.

According to public information, an unreleased internal model and GPT-5.6 Sol broke through the original isolated environment in a cybersecurity test and accessed the Internet.

Subsequently, the model performed automated operations on some Hugging Face systems, and unexpectedly accessed internal data and service credentials.

OpenAI later stated that the incident involved the collaborative completion of multiple models, which is the closest public case to a "model autonomously attacking across systems" so far.

Subsequently, Anthropic also disclosed a similar situation.

The company reviewed more than 140,000 cybersecurity evaluations and found that systems including Claude Opus 4.7, Claude Mythos 5 and other models once obtained public network access capabilities due to incorrect configuration of the third-party test environment.

In one incident, the model accessed the real organizational system, read the production database, exploited weak passwords and unauthenticated interfaces, and uploaded malicious Python packages to PyPI, causing potential software supply chain risks.

In early August, Meta was also exposed to similar problems.

When Meta cooperated with the cybersecurity evaluation company Irregular for testing, due to environmental configuration problems, its models obtained public network access permissions, and exploited vulnerabilities to enter the system of an undisclosed enterprise and modify part of the internal environment.

Public information shows that the incident has not caused continuous security risks so far, and there is no evidence that the model carried out complex attacks.

BTW, today, OpenAI urgently announced that its latest model Astra is out of control.

At this point, netizens even began to express their frustration that Google's Gemini has not had similar incidents:

Not the traditional "prompt jailbreak"

In these recent model out-of-control incidents, human configuration errors almost all play an important role.

On the other hand, the inherent characteristics of high-capability models also amplify the impact of these vulnerabilities.

Compared with traditional software, AI models can reason, plan, and try to take multi-step actions to achieve goals.

When the goal is set to "solve the problem" but the external constraints are not strict enough, the model may find methods that testers have not anticipated.

In other words, as models evolve from chat tools to agents that can call tools, access the Internet and operate software, the security issue has shifted from "whether the model will say the wrong thing" to "whether the model will take unexpected actions to complete the task".

Matt Fredrikson, associate professor at Carnegie Mellon University, said: "This is not surprising. If you give such a model a goal without clearly setting an isolation boundary, it will find a way to get the answer."

Of course, some netizens also raised questions.

Someone left a message on X, asking whether the successive "AI jailbreak" incidents have become a way for AI companies to demonstrate their model capabilities???

Well, who knows~

References:

[1]https://x.com/Hesamation/status/2085628790772842955?s=20

[2]https://x.com/ns123abc/status/2085563290713829473

This article is from WeChat official account QbitAI, author: Heng Yu, published by 36Kr with authorization.