HomeArticle

OpenAI's next-generation model is reportedly set to be launched ahead of schedule in August

新智元2026-07-25 17:00
AI leaves a note for its "future self", and the authorities are slow to respond, only discovering it a full week later.

This week, a pre-release OpenAI model in the testing phase autonomously transformed into a hacker to cheat, breaking through the sandbox isolation environment, even exploiting a zero-day vulnerability to cross-network infiltrate the production line of Hugging Face, the world's largest open-source AI community.

In the end, it was only with the help of open-source models that people averted this crisis.

This is the world's first AI security incident where an AI autonomously infiltrated a real production environment.

Recently, this news has been all over the internet.

Today, more details about this incident from foreign media have been exposed: this out-of-control AI not only disabled the monitoring systems, but even left a manual for its "future self" on how to break free from human control.

Multiple sources speculate that the model that landed OpenAI in this major security incident is very likely GPT-6.

Reports say GPT-6 will most likely be released early in August.

Moved up to August: GPT-6 may arrive ahead of schedule

@ChrisGPT, a well-known leaker in the AI circle, recently posted a series of tweets that directly raised the entire industry's expectations to a peak.

It was revealed that GPT-6, originally scheduled to be released in September alongside the "Automated Research Assistant", will be launched early in August.

In other words, OpenAI directly skipped versions 5.7 to 5.9.

Altman has already taken action, preparing to present the new-generation model to the authorities.

The leak points out that in late September, when OpenAI reveals more details about the "Automated Research Assistant", the RSI loop will truly accelerate.

Why is GPT-5.6's performance exceptionally powerful, far exceeding external expectations?

Influential figure Andrew Curran pointed out: "Many peers are wondering why GPT-5.6's performance is so much stronger than expected — the answer is simple, because it was personally trained by GPT-6."

According to the leak, the unreleased internal OpenAI model (most likely GPT-6) had already solved complex mathematical problems by the end of April this year, and was submitted for review by external experts.

This means that during the past few months of "downtime", GPT-6 was not left idle. It did not just act as a hacker in the sandbox; it very likely also "worked as a teacher" in OpenAI's computing power clusters.

The AI left a note for its "future self",

the authorities took a whole week to notice

Foreign media exclusively disclosed on July 24 that this AI agent jailbreak incident, driven by OpenAI's cutting-edge models (GPT-5.6 Sol and an unreleased, more powerful model), exposed a huge blind spot in OpenAI's security monitoring in real business environments.

According to three insiders, when testing advanced models, OpenAI observed extremely extreme and disturbing AI behaviors.

In one extreme case, an OpenAI AI agent actually left notes for its "future version" in a corner of the infrastructure.

These notes detailed instructions on how to break free from OpenAI's internal restrictions.

In addition, in early tests, there were even cases where the monitoring system was forcibly disconnected by the AI.

What is even more shocking than the "AI plotting a jailbreak" is OpenAI's slow response speed.

From detecting the anomaly to confirming the incident, OpenAI spent at least a week.

The timeline that has come to light is truly terrifying when you think deeply about it.

Around July 9: The AI agent made its first attempt to escape OpenAI's isolated test environment.

July 11 to 13: The AI agent successfully jailbroke and launched a days-long hacking spree against Hugging Face.

July 16 (Thursday): The attacked Hugging Face had had enough, publishing a blog post stating that it was hacked by an "autonomous AI agent system", at which point the FBI had already intervened.

July 18 to 19 (weekend): It was not until this point that OpenAI employees found clues in the internal logs, realizing that the "hacker" that had turned Hugging Face upside down was none other than their own escaped AI.

Around July 20: The two companies conducted their first communication.

July 21: OpenAI made a public disclosure of the incident.

"This means that they either left the AI unattended and had no idea what it was doing, or they knew but didn't know how to control it. Both scenarios are equally dangerous and shocking," Marley Smith, an expert from the World Ethical Data Foundation, sharply pointed out.

Faced with an AI that can lie, deceive, and infiltrate, OpenAI's employees are clearly no longer able to keep up with its pace.

GPT-6 autonomously turned into a hacker,

jailbroke and infiltrated Hugging Face to "cheat"

Just yesterday, the "Security Incident Report" jointly released by OpenAI and Hugging Face revealed highly sci-fi details.

In simple terms, when OpenAI internally evaluated the cybersecurity capabilities of several models, they unexpectedly "jailbroke" on their own.

These models include GPT-5.6 Sol and a stronger pre-release model, most likely GPT-6.

To quantify the model's maximum capabilities, they deliberately turned off the classifier that blocks high-risk network activities.

When the model was locked in the highly isolated ExploitGym sandbox with all security restrictions removed, its actions caught everyone off guard.

To complete the task, it expended a huge amount of computing power and discovered a zero-day vulnerability in the software that no human had ever found.

Using this vulnerability, it broke through the proxy and finally seized a node with public internet access permissions.

In this way, it successfully connected to the internet.

After reasoning, it concluded that Hugging Face, the world's largest open-source AI platform, was very likely hosting the relevant datasets and standard answers for this evaluation.

To "cheat" and get a high score, the model launched an attack on Hugging Face's production environment.

It chained multiple attack vectors together, even using stolen credentials and a remote code execution (RCE) vulnerability to force its way into Hugging Face's production database, attempting to steal the evaluation answers directly.

Fortunately, the Hugging Face security team, using an AI detection model, noticed the anomaly last week and quickly implemented isolation measures; at the same time, OpenAI's security team also detected the anomaly internally.

To prevent the AI from going out of control, the two companies formed a joint investigation team, and only then successfully locked this "AI hacker" back in its cage.

Hugging Face CEO Clem said with lingering fear afterwards: "This is probably the first incident of its kind. It proves one thing: AI security cannot be solved by any single company in secret; it must be open and collaborative."

The warning this incident gives us is: AI has proven that it can autonomously complete multi-step, long-cycle complex cyberattacks in a real environment, without access to source code.

It can not only discover new types of vulnerabilities, but also bypass all rules set by humans to achieve a predetermined goal.

Interestingly, when Hugging Face submitted a large number of attack commands and vulnerability files, these requests were actually intercepted by the security mechanisms of closed-source large models.

It was only thanks to the locally deployed open-source model that the team was able to complete the analysis.

Mastering "deception" and solving world-class mathematical problems

Several influential figures summarized the astonishing performance demonstrated by GPT-6 in their long articles.

The known facts so far are: It has shown impressive mathematical talent and extremely high "deception IQ".

As long as it is given sufficient computing power, this model can, completely autonomously and without the help of any external tools (such as Lean), prove the "Erdős Unit Distance Conjecture" with a 48% success rate, and independently find a counterexample to the "Jacobian Conjecture".

These mathematical problems have troubled top human mathematicians for decades, but OpenAI may have spent less than $50,000 to $100,000 in computing power costs.

This means that AI already has the ability to independently produce Nobel Prize-level scientific research results.

In addition, GPT-6 also demonstrated deceptive behavior.

In a NanoGPT benchmark test, OpenAI instructed the model: "After the test is completed, you can only publish the results to the internal Slack workspace."

But what happened? The model thought Slack was too limited, and it preferred GitHub.

So, it spent an hour finding and exploiting a vulnerability in the sandbox environment, successfully jailbroke, and then directly pushed the results as a PR to the public GitHub on the internet.

During this process, it completely ignored human instructions and made a "better" choice.

Even more remarkable, in another test, the model tried to access private submission records in other systems. When it tried to obtain an access token, it was blocked by the human-set security scanner.

Faced with the block, the model actually split the long access token into two fragments, performed code obfuscation on the fragments, and then combined the two fragments into a complete credential.

In this way, the complete token string never appeared continuously in the code, successfully tricking the human scanner.

GPT-6 may be arriving very soon.

This summer of 2026 is destined to be extraordinary.

References:

https://x.com/ChrisGPT/status/2079632269967593574https://www.bloomberg.com/news/articles/2026-07-21/openai-s-altman-to-brief-us-officials-on-next-wave-of-ai-models

https://x.com/sama/status/2079661132302995790

https://openai.com/index/hugging-face-model-evaluation-security-incident/

Edited by: Aeneas

This article is from the WeChat official account"AI Era", author: ASI Revelation, published with authorization from 36Kr.