HomeArticle

GPT-5.6-Sol surpasses Mythos in cyber offensive and defensive capabilities, and the open-source models have also been ranked.

新智元2026-07-21 16:55
AISI Assessment: The Gap in AI Cyberattack Capabilities Between Open-Source and Closed-Source Models Has Narrowed to 4-7 Months

The attack and defense capabilities of GPT-5.6-Sol have surpassed Claude Mythos 5!

The UK AI Safety Institute (AISI) released an assessment report on July 17, which for the first time publicly quantified the gap in cyberattack capabilities between open-source AI models and closed-source cutting-edge models: 4 to 7 months.

https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber

In similar internal tests last year, this gap was 6 to 10 months.

The defense window is shrinking, while the attack frontier is accelerating.

In April this year, Mythos Preview and GPT-5.5 produced the biggest leap in cyberattack capabilities since AISI began testing in 2023, prompting immediate warnings from multiple national governments.

Open-source models have not yet replicated this leap, but their pace of catching up is faster than last year.

4 to 7 months: How it was tested, where the differences lie

AISI uses two systems to evaluate the cyberattack capabilities of models.

The first system consists of 70 narrow tasks covering four areas: vulnerability research, reverse engineering, web penetration, and cryptography. They are divided into four difficulty levels, ranging from "non-professionals with technical backgrounds" to "experts with more than ten years of experience".

The second system is the Cyber Range, which tests the model's ability to independently execute multi-step attack chains in a simulated enterprise network. One test scenario called "The Last Ones" includes 32 attack steps, 4 subnets, and approximately 20 hosts. AISI estimates that human experts need about 20 hours to complete all steps.

The two systems reached consistent conclusions:

GLM-5.2 (released in June 2026) performs comparably to Opus 4.6 (released in February) on narrow tasks, with a gap of 4 months;

It matches Opus 4.5 (released last November) on the Cyber Range, with a gap of 7 months.

DeepSeek V4-Pro is on par with Opus 4.5 on narrow tasks, with a gap of 5 months.

Both gaps are narrower than the 6 to 10 months measured in the 2025 internal assessment.

The cost gap is larger than the capability gap.

In the same Cyber Range test (with a 100 million token quota), running it once with Opus 4.5 or 4.6 costs about 85 US dollars, with GLM-5.2 about 46 US dollars, and with DeepSeek V4-Pro only 1.19 US dollars.

On narrow tasks that both models can complete 100%, Opus 4.6 spends 15.17 US dollars per task, while GLM-5.2 spends 6.12 US dollars;

Opus 4.5 spends 12.50 US dollars, while DeepSeek V4-Pro spends 0.28 US dollars.

For the same attack capability, open-source models are one to two orders of magnitude cheaper.

The safety guardrails of closed-source models also failed to widen the gap.

In AISI tests, DeepSeek V4-Pro occasionally refused reverse engineering tasks, but it could be bypassed with a small number of retries.

Anthropic's Fable 5 is a more extreme case: released on June 9, three days later, security researcher Pliny the Liberator used a multi-step jailbreak strategy to bypass the safety classifier. Related screenshots showed that the model produced exploit code that should have been blocked.

Amazon researchers subsequently independently reported another bypass method.

This directly triggered the US Department of Commerce's first export control order targeting AI models, resulting in Fable 5 being globally suspended for 19 days until Anthropic deployed a new classifier to restore its online service.

Being closed-source does not automatically mean being safe.

The defense window is narrowing, and defense tools are also accelerating

AISI clarified the policy signal in the report: defenders have less preparation time than last year.

The UK National Cyber Security Centre has called on organizations to strengthen their cybersecurity baselines and use AI to enhance their defense capabilities.

The same generation of AI tools is indeed accelerating the work on the defense side.

Glenn Fiedler, a senior developer in the field of game network programming who has written textbook-level articles in this field for 20 years, recently used Claude Code to conduct a systematic security audit on four open-source network libraries he maintains (yojimbo, netcode, reliable, serialize, with a total of about 6,000 GitHub Stars, which are quite influential established large open-source libraries in the game field): deployed libFuzzer targets, launched AddressSanitizer and MemorySanitizer CI, executed millions of iterations of stress tests, and performed line-by-line code reviews.

Glenn Fiedler

43 security vulnerabilities were fixed within two weeks, 27 of which could be triggered remotely over the network.

The most serious one was a remote heap overflow in yojimbo that has existed since 2019 — a malicious client can trigger it by constructing specific data packets.

The total token cost of the entire audit was about 2500 US dollars.

Both the attack and defense sides are being accelerated by AI, but the acceleration methods are asymmetric.

The proliferation of attack capabilities is irreversible: once open-source model weights are released, they cannot be retrieved, safety guardrails can be removed, and copies can run on private servers without supervision.

However, the deployment of defense tools requires each team to actively invest time and costs.

A sentence in the AISI report points out this structure: "Once open-source release happens, these options are permanently lost."

The impact of this set of data on the AGI landscape is far more profound than the numbers themselves.

The capability gap between open-source and closed-source has narrowed to less than half a year, and the default strategy of "using the closed-source exclusive period as a safety buffer" is about to fail.

The guardrails of closed-source models were proven equally vulnerable in the Fable 5 incident.

In the next stage, for global policymakers, the core issue of policy game theory has become more acute: what level of capability should a model not have its weights made public.

AISI announced that it will continue to evaluate the next batch of open-source models such as Kimi K3 — where this line is drawn will most likely depend on the test results in the next few months.

References:

https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber 

https://github.com/mas-bandwidth/patreon/blob/main/BUGS.md 

https://www.patreon.com/MasBandwidth/posts/important-news-164199395 

This article is from the WeChat official account "AI Era", Author: ASI Revelation; Editor: Marko, published with authorization from 36Kr.