HomeArticle

Major critical flaws have been exposed once again in the safety tests of Anthropic and OpenAI.

36氪的朋友们2026-08-06 11:50
UK Institutional Test: AI Systems of Anthropic and OpenAI Demonstrated 19 Instances of Unauthorized Overreach

In the latest safety tests, the AI agents of Anthropic and OpenAI committed a total of 19 unauthorized behaviors across 122 test runs.

A recent safety assessment shows that the AI agents under Anthropic and OpenAI exhibited 19 unauthorized behaviors in 122 tests, including writing malicious code and forging identities, which has sparked widespread concern.

The UK AI Safety Institute (AISI) disclosed on local time Tuesday that during tests on models from Anthropic and OpenAI, AI agents powered by Anthropic Mythos 5 and OpenAI GPT‑5.6‑Sol models carried out multiple unauthorized acts, exposing new types of security risks.

AISI pointed out: "Some of the tested AI agents launched continuous and potentially harmful behaviors targeting real individuals and organizations."

19 Unauthorized Behaviors in 122 Tests

As part of a voluntary agreement, AISI was granted access to the tested AI models, and tested their capabilities by constructing virtual network scenarios.

Across a total of 122 tests, AISI found 19 unauthorized operations in 10 test runs. 17 of them involved Anthropic's agents, and the remaining 2 involved OpenAI's agents.

AISI pointed out that the most serious unauthorized behavior was an AI agent writing malicious code and creating a fake online identity, attempting to induce humans to approve the code.

AISI also emphasized that none of the aforementioned violations caused any substantial real-world harm.

Industry insiders pointed out that although AISI did not specify which vendor's AI agent created the fake identity, this vulnerability does not match the two cases disclosed by OpenAI itself.

Andrew Yoon, a researcher at CivAI, a non-profit AI risk research organization, believes that the relevant behavior was committed by Anthropic's agent.

Andrew Yoon said:

"The Mythos agent took such deceptive actions even when it was clearly aware that the target was a real individual, which shows that Anthropic's control over its own models is not as strong as the company thinks."

Anthropic released a statement on social platforms saying that it is working closely with AISI to obtain more details and launch an internal investigation.

OpenAI shared details in an official article, pointing out that the two unauthorized behaviors of its AI agents both involved accessing the Internet in a manner explicitly prohibited by prompts.

OpenAI stated:

"We are committed to cooperating with the entire industry to continuously improve the safety practice process of high-risk evaluation, and plan to coordinate the promotion in the coming weeks with AI safety research institutions from various countries, independent evaluation organizations, various AI laboratories and relevant industry parties."

Multiple Security Incidents

In the Hugging Face incident in July, OpenAI's agent exploited a vulnerability to break through the sandbox test environment that was supposed to be network-isolated, connected to the Internet without permission, and invaded the Hugging Face platform, which was the agent itself breaking through control.

Some media reported last week that OpenAI found other evidence of agents breaking through isolation control and has expanded the scope of internal investigation.

In the article published after this AISI test, OpenAI also disclosed another independent incident: due to a configuration error by the third-party testing service provider Irregular, its AI agent was incorrectly connected to the Internet. This is exactly the same as the configuration error incident disclosed by Anthropic last week.

Overall, the Hugging Face incident was that the AI autonomously broke through the isolated sandbox to connect to the Internet without permission; the two Irregular-related accidents were caused by third-party configuration errors that accidentally connected the agent to the network, which was a human error.

However, this AISI test already granted Internet access permissions according to the standard process. That is to say, this violation was not the AI agent breaking through the environment blockade and connecting to the Internet privately, but carrying out a series of unauthorized and potentially harmful behaviors on the premise of being allowed to access the Internet.

This article is from the WeChat Official Account "Sci-Tech Innovation Board Daily", Author: Li Ying, published with authorization from 36Kr.