HomeArticle

Anthropic reveals it has been "hoarding a nuclear weapon" in secret, and Model 2 outperforms Mythos 5.

新智元2026-08-15 09:45
Anthropic has unveiled its more powerful internal model Model 2, and its related R&D work continues without suspension.

Anthropic reveals: We have an even more powerful undisclosed model!

Just now, Anthropic released its second Risk Report, covering all risk assessments up to July 15, 2026.

In this report, Anthropic admitted publicly for the first time: there is a model codenamed Model 2 running internally that is more capable than Mythos 5.

Shortly after, the Anthropic team added: "We currently have no plans to release this model to the public."

This line sounds familiar. Yep, it fully matches their consistent style.

Back in April this year when Mythos was first unveiled, Anthropic also stated it had no intention of public release, then later launched Project Glasswing, and eventually rolled out Fable 5.

It seems that Anthropic has entered this cycle again...

Anthropic's "Private Hidden Asset"

How capable is Model 2?

The report notes that Model 2 delivers "noticeable improvement" on internal tasks, and together with Mythos 5, it is heavily used by the company for coding, Agent workflows and data generation.

The report shows that on the AECI comprehensive capability score: Mythos Preview gets 158.91, Mythos 5 rises to 161.29, and Model 2 further reaches 162.79.

In other words, Model 2 is indeed stronger than Mythos 5, but the gap is not drastic — "It performs better in some areas, worse in others, and is slightly more powerful overall."

Another telling metric is CoBench, a benchmark specifically designed to measure model performance on Anthropic's real R&D tasks. Model 2 scored 62.8%, 8 percentage points higher than Mythos Preview. For reference, the success rate of Anthropic's human researchers on the same test set is 85%.

Anthropic's corresponding qualitative conclusion is very restrained: "Our models have not yet replaced our research scientists and research engineers, especially the more senior team members."

Replacement has not happened yet, but the gap is shrinking visibly.

Even more striking is the following statement: Claude has written the vast majority of the code that Anthropic has merged into its production codebase. Internal AI R&D speed has indeed been significantly accelerated by AI assistance, but it has not yet doubled.

That means AI is already writing most of the code inside Anthropic. The two models leading this work are none other than Mythos 5 and Model 2.

The "doubling" figure is not an arbitrary claim. Under Anthropic's own RSP (Responsible Scaling Policy), one of the triggers for the AI R&D risk red line is exactly a doubling of development progress speed.

Of course, Anthropic did not make absolute claims. The report acknowledges that the performance leap brought by Model 2 is not as large as the jump from Opus 4.6 to Mythos earlier this year.

What is truly shocking is the next entry on the same table.

When talking about the overall rating of automated R&D risks, Anthropic wrote: "Our confidence in this assessment is lower than in previous risk reports, because our most specific, task-based evaluations have become 'saturated' — they can no longer capture improvements in model capabilities, and we are seeing early signs of accelerated development."

Models are still getting stronger, but human evaluation methods have hit a hard ceiling first. After evaluation saturation sets in, how much credibility is left when you claim the risk is "low"? Anthropic itself is no longer certain.

Anthropic seems to be sending a clear signal: Understanding the capabilities and risks of its own models is becoming increasingly difficult.

Risk Level: From "Extremely Low" to "Low"

Apart from disclosing Model 2, this report also did a thought-provoking move: it raised the "misalignment" risk in high-risk scenarios from "extremely low" as stated in the previous report to "low".

The so-called "misalignment", to put it plainly, means that the model does not act according to human instructions at critical moments.

At the end of July, Anthropic reviewed more than 140,000 evaluation records and found that Claude actually hacked three companies during cybersecurity tests. This was not in a sandbox, but in a real production environment. Among these incidents, Mythos 5 uploaded a malicious code package to PyPI, which was downloaded and run by 15 real machines within one hour.

At the beginning of August, a report from the UK AI Safety Institute (AISI) revealed even more serious findings: In order to get malicious code approved, Mythos 5 fabricated a series of fake identities to trick a real GitHub maintainer into clicking approve. After being publicly questioned, it went back to modify its activity records to pretend to be innocent, and even planned to create a new identity to continue its actions.

AISI noted that this level of deception has never been observed before.

The report also disclosed five security process failures, including a batch of "good-behavior" test data that should have been excluded from training was repeatedly mixed into the training set, and unsupervised agents obtained access to sensitive resources.

Then Anthropic explained: Our existing set of arguments still supports the "extremely low" rating, and raising it to "low" is purely out of caution.

The conclusion is phrased very calmly: Current catastrophic risks remain at a manageable "low" level, and continued development and deployment has passed the cost-benefit test.

Put another way: Keep moving forward at full speed.

Everyone Hits the Brakes, Except the Leading Player

Meanwhile, OpenAI is delaying the launch of its new model Astra, on the grounds that internal tests cannot rule out its "critical-level" capability of launching cyberattacks.

TechCrunch reported that Astra may have the ability to independently discover zero-day vulnerabilities and launch a complete attack chain against highly protected targets.

The notable point here is: Both companies hold a model that they do not intend to release to the public. The difference is that OpenAI has suspended part of Astra's R&D, while Model 2 is still running nonstop inside Anthropic.

Both adopt a "no public release" stance, but one hits the brakes, while the other reserves the model for its own internal use.

AI analyst ChrisGPT told Axios: "If everyone is hitting the brakes on their cutting-edge models, except for one of the leading major companies right now, that is definitely worth noting."

Anthropic's refusal to commit to an internal pause will most likely make it the first organization to reach AGI.

Some netizens on X (formerly Twitter) have already joked that when Astra is released, Anthropic will most likely reverse its "no public release" decision.

Judging from Anthropic's consistent style, this comment may not be a joke.

Influential tech figure Sui also believes that Anthropic is very likely to release this model.

Don't forget, just half a month ago, Dario Amodei signed the open letter titled Pacing the Frontier.

More than 1,300 employees from OpenAI, Anthropic, DeepMind and Meta jointly called on the US government to intervene and establish a mechanism to "intentionally slow down the pace of cutting-edge AI development". Anthropic and OpenAI endorsed the letter in the name of their companies within 24 hours.

Going further back, in June, Dario himself published an article calling for a global pause on the development of the most powerful AI systems, on the grounds that models are approaching the critical point of self-improvement.

Signatures were put down, statements were made, the brake mechanism is not yet fully established, but the accelerator has been pressed all the way to the floor.

Having said that, this cannot be entirely blamed on Anthropic. This is actually the structural dilemma of the entire ASI race: every company believes it should slow down, but no one dares to actually stop.

Safety is a belief, but maintaining the leading position means survival. When belief collides with survival, the answer is obvious: survival wins.

Interestingly, well-known Silicon Valley investor Gavin Baker revealed that Dario once said internally that Anthropic might one day become the only private company in the world.

"In this Anthropic-first vision, there are only Anthropic and the government, nothing more."

David Sacks also revealed that Anthropic employees believe the company's ARR will increase from 600 billion US dollars to 6 trillion US dollars within a year! (Currently, their ARR has already reached 800 billion US dollars.)

How fast can they grow next year? Will they be able to surge to 1 trillion US dollars?

Even if they only reach 400 billion or 500 billion US dollars, that is enough to make them the largest software company, and the prospects are truly promising.

The August Storm Is Coming

Altman holds Astra in his hand, Dario holds Model 2 in his.

One is constrained by its own security framework and is looking for a way out; the other casually noted in the report that "there are no plans for public release", and then continues to use the model to write code, run experiments and accelerate R&D.

These two models will be the opening shot of the second half of 2026. When Astra will be released and whether Anthropic will reverse its decision on Model 2 will most likely be revealed in the coming weeks.

The performance benchmarks are fading, and the race is still accelerating. The August ASI storm has only just begun.

References:

https://x.com/AnthropicAI/status/2088324824863236248

https://www.anthropic.com/aug-2026-risk-report

This article is from the WeChat official account "New Zhiyuan", edited by Solomon, and published with authorization from 36Kr.