HomeArticle

Breaking | The most powerful ChatGPT model hits emergency brakes, Sam Altman: You (Astra) scare me.

爱范儿2026-08-08 09:21
OMG, you (Astra) scared me.

OpenAI, which has always been adhering to the principle of "push as hard as you can until you cannot move forward", has surprisingly hit the brakes on its own initiative.

And this is not a minor tweak, but a direct slowdown of its R&D progress. 

The whole story starts from Astra, an unreleased next-generation large model that OpenAI has been intensively developing internally. 

Official blog 🔗 https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/ 

After multiple rounds of internal assessments, OpenAI found that the model's capabilities in cybersecurity may be so strong that it can no longer be treated as an ordinary AI product. 

Writing code, identifying vulnerabilities, automatically planning attack steps, invoking tools, continuously executing tasks... 

Once these capabilities are combined, Astra is very likely to become a "stat monster" with a full skill tree of hacking capabilities. 

OpenAI even stated bluntly: "We cannot rule out the risk that Astra possesses critical cyberattack capabilities." Wow, the cyber R&D process has accidentally produced a rogue AI that breaks all preset rules! 

To prevent potential risks in advance, OpenAI now has to roll out its ultimate measure. 

They will comprehensively expand testing and security protection for Astra, and at the same time actively slow down its R&D speed. The R&D pace will be restricted until new security protection measures are fully in place. 

However, OpenAI obviously does not plan to keep it locked in the lab forever. Sam Altman later personally came out to explain that Astra is a "powerful model", and OpenAI still hopes to eventually make it accessible to all users.

In his view, keeping the most powerful models in the hands of a small number of people for a long time is not a good strategy. But considering the demonstrated cyber capabilities of Astra, it is expected to take "a little more time" to resolve the security issues.

This strict standard is based on the "Preparedness framework" that OpenAI first released back in 2023. 

It is a typical case of lifting a rock only to drop it on one's own feet: the rules they set for themselves have to be followed even if it causes inconvenience. 

At the recently concluded Black Hat cybersecurity conference, OpenAI's technician Michael Dalton also publicly acknowledged that OpenAI has begun to "consciously slow down research to enhance safety".

To place Astra in a sufficiently controlled environment, OpenAI directly arranged for Astra to get the "VIP single room" treatment. 

These measures include but are not limited to: a completely isolated test environment, and full-coverage universal monitoring for all Agentic applications without any blind spots. 

What is even more rare is that Sam Altman is extremely cautious this time with a strong sense of risk prevention. 

According to sources from White House officials in the United States, OpenAI "voluntarily" notified the US government of its plan to delay the release.

This is probably the first time a cutting-edge AI laboratory has voluntarily promised to slow down the progress of its core model out of fear of the destructive power of the model in the cyber world.

After all, in this era dominated by AI FOMO (Fear Of Missing Out), everyone is frantically scaling up their models, and no one dares to stop easily. 

Speaking of which, we have to mention the "brilliant operation" of its peer Anthropic. 

Their past track record can be regarded as a textbook example of repeated flip-flopping. Anthropic once set a lofty goal that if the terrifying capabilities of AI exceed human control, training must be suspended. 

However, in February this year, when updating its "Responsible Scaling Policy", they quietly removed this suspension clause. 

Their excuse sounds perfectly justifiable: if only one AI developer pauses to focus on security while other peers are frantically deploying models without adequate safeguards, the world will become even more unsafe.

Anthropic is like: I don't want to compete blindly, I'm forced to do this, folks! (just kidding) 

In contrast, Sam Altman, who is always good at hype, even looks quite honest and reasonable now.

But you have to pay for what you do sooner or later. 

In June this year, Anthropic still released their model with the strongest cyberattack capabilities as usual, that is, Fable 5, the security-stripped version of Mythos 5. 

Not only that, in an official blog post in June, they even publicly warned: AI is showing the ability of self-evolution! And called on the whole world to pause AI development together. 

Wow, it turns out that everyone is competing fiercely in the AI track, but you run out to act as the savior. 

Looking at the broader environment, the background of this wave of incidents is even more thought-provoking. 

At present, the US government is stepping up efforts to establish an assessment framework for AI models before their release. 

This week, selected leading enterprises in the industry just attended a closed-door briefing on this framework. 

However, many questions remained after the meeting. For example: How should enterprises carry out daily docking with the government? How long will this official review process take? What exactly do the government and the industry want to learn from each other?

What's worse, who will own these models and who will review them has not been determined yet. 

Although the framework talks a lot about "national-level risks" and "state-of-the-art models", it does not give clear definitions at all. It's just like listening to a speech that says nothing new (thinking.jpg). 

Interestingly, in the coverage of OpenAI's incident, the most notable point is actually a lightly worded sentence specially added by the official: "Astra has nothing to do with the previous Hugging Face vulnerability incident."

Wait a minute... the first law of AI is that only what the manufacturer officially denies is credible. 

Could it be that this model also created a message board in the shared software package manager inside OpenAI, and quietly passed small notes out through the network cable several times? 

This article is from the WeChat official account "APPSO", the author is the team that discovers upcoming products, and 36Kr publishes it with authorization.