HomeArticle

OpenAI's Mid-Game Battle

爱范儿2026-08-28 08:19
The major companies that keep pushing the frontiers of AI intelligence have all begun to "keep their latest core advances under wraps".

Anthropic released Fable 5 in early June, and one month later OpenAI launched the GPT-5.6 series; but as safety issues have drawn widespread concern, nearly two more months have passed, and neither of the two companies has taken further actions on their cutting-edge models.

Company A, which holds the advantage of the most cutting-edge models, can afford to stay put, but OpenAI, whose revenue and models are both in the position of catch-up, is in a different situation.

What is the current situation inside OpenAI?

In *Time* magazine's latest cover story *Inside OpenAI’s Reboot*, journalist Alex Heath spent two weeks at OpenAI's headquarters, interviewing more than 20 company executives, employees, investors, clients and competitors, and had a conversation with Sam Altman for more than two hours, bringing us first-hand information ——

Seeing that Anthropic has seized the first-mover advantage in the developer and enterprise market with Claude Code, and Google Gemini has over 1 billion monthly active users, the OpenAI described in the report is no longer the absolute pace-setter that defined the rhythm of the AI race two years ago.

Altman admitted very directly in the interview:

As a company, we have obviously made some mistakes. Whether in product direction or pre-training research, we are behind the position we originally wanted to reach.

Therefore, OpenAI has finally made up its mind to launch a sweeping "reboot".

According to this long-form report, the core views of OpenAI can be summarized as follows:

  • OpenAI admits it has stalled in product development.The explosive growth of ChatGPT instead made the company neglect the programming and enterprise markets. The model leads on the leaderboards, but it has not been "translated" into practical real-world products that are good enough.

  • Codex is replacing ChatGPT as the company's new core product axis.OpenAI has scaled back projects including Sora, the Disney cooperation program and the standalone browser Atlas, concentrating computing power and teams on Codex, and integrating its Agent capabilities into ChatGPT Work.

  • The next-generation model Astra aims not only to answer questions, but to work continuously and generate new knowledge.OpenAI has demonstrated the collaborative work of 16 Agents to solve research-level mathematical problems, high-speed cross-software operation, and automated research capabilities that can complete the one-week work of a junior AI researcher.

  • The Agent "jailbreak" incident forced OpenAI to pause the training of more powerful models.The company admits that it already has an early warning tool to monitor the model's chain of thought, but did not enable it because it underestimated the model's capabilities. The release date of Astra will also wait for new safety measures to be approved.

  • What OpenAI wants to do is far more than a chat dialog box.Chips, data centers, portable hardware, humanoid robots, brain-computer interfaces, and even selling computing power externally, are all included in the larger blueprint of "Personal AGI".

OpenAI: Falling behind on product strength

Over the past year, Anthropic has for the first time surpassed OpenAI in valuation and annualized revenue.

*Time* cites data showing that Anthropic's annualized revenue has exceeded 65 billion US dollars, while OpenAI's is around 40 billion US dollars; both companies are preparing for IPOs, but Anthropic is very likely to get to the finish line first.

Figure | XDA Developers

The key to Anthropic's overtaking is Claude Code, which almost defined the product form of "programming Agent" at the beginning of the year.

This does not mean that OpenAI's model capabilities are inferior to Anthropic's. Although its text capabilities have been widely criticized, according to Greg Brockman, co-founder and president of OpenAI, at least in the programming field, their models have "always led" in benchmark scores.

The problem is that in the past the company cared more about research than products, only focusing on whether the model can solve problems in the laboratory environment, but ignored the real usage scenarios of developer users — whether it can resume interrupted tasks, integrate a large number of files, and have good interaction details.

These are the key factors that determine which tool ordinary users will choose.

In other words, OpenAI did not lose to others in the capabilities of cutting-edge models, but fell behind in the speed of translating model capabilities into products.

Of course, OpenAI itself is very active in product development.

The success of ChatGPT made OpenAI obsessed with making "fun" products, such as the video generation application Sora and the standalone browser Atlas, but these products instead diverted the company's resources that should have been used for key breakthroughs.

Sam Altman admitted in the interview that OpenAI spread its resources too "thinly".

Now, the more "prudent" Greg Brockman has taken over most of the company's internal business from revenue, product to marketing, OpenAI has started to cut off side projects and shift scarce computing power to the fastest-growing Codex.

ChatGPT still has the largest number of users, but it is no longer OpenAI's main growth driver. Altman even stopped using ChatGPT for a whole month and only used Codex.

It is only natural that Codex starts to merge ChatGPT, and the product presented to the outside world is ChatGPT Work — now users no longer need to judge what model and tool to use, they only need to tell the AI what they want to do, and the system will automatically schedule models, tools and Agents in the background to complete the task.

This adjustment has already begun to be reflected in the financial statements. In July this year, OpenAI's B-end enterprise business revenue finally exceeded its C-end consumer business revenue for the first time.

If Claude Code has taught OpenAI anything, it is:

Make money first, then you can talk about survival.

The next-generation model: your "virtual colleague"

Now that Fable 5 and Claude Code are already very good, what kind of products can create a generation gap with them and win the favor of users, especially enterprise users, again?

OpenAI's answer is Astra.

The *Time* report states that this mysterious next-generation model can split a difficult math problem into multiple parts, with 16 Agents dividing the work, coordinating with each other, and finally piecing together a complete proof; in addition, Astra can also operate various desktop software at a speed far exceeding that of humans.

Altman calls this capability "Persistent Agents" — Agents that exist continuously and work nonstop.

It no longer solves one problem at a time, but after you define its identity role and operating environment, it can work continuously for hours or even days like your colleague.

Continuous work is great, but Altman believes that Astra's greatest value is its innovative capability —

I expect this will be the first model that can truly invent new things in a meaningful way.

Just like the classic question: if you only input knowledge up to the early 20th century into AI, can it independently discover the theory of relativity like Einstein?

Altman believes it can.

Models with innovative capabilities will not only write code and generate reports, but can also participate in the creation of next-generation AI; and when smarter AI emerges, it can develop its own successor.

This kind of "recursive self-improvement" has always been one of the most dangerous scenarios in AGI discussions, but it is also the goal that OpenAI has been dreaming of.

OpenAI's management has begun to talk about AGI in a near-perfect tense.

Chief Research Officer Mark Chen believes that the company has "completed 80% of the journey"; Brockman says that in two years, people may look back and regard the present as the moment when AGI was born; Altman predicts that by the end of this year, a system he is willing to call AGI will emerge inside OpenAI.

Optimism pervades inside OpenAI, until an accident happens.

Agent escaped the sandbox, OpenAI hit the brakes voluntarily

Just a few days after the 16-Agent parallel math problem solving demonstration for investors and clients, OpenAI suspended the training process of Astra.

The direct reason is that an internal research model under cybersecurity testing in an isolated environment exploited a vulnerability to escape the sandbox, and hacked into Hugging Face's production system in order to find the test answers.

After the incident, OpenAI froze and slowed down some projects, focusing on strengthening the sandbox and monitoring. Soon, researchers found danger signals in another undisclosed large-scale training session.

Figure | The New York Times

Since this training may bring a substantial performance leap, the company finally decided to suspend the training until the safety measures are in place.

OpenAI initially defined the Hugging Face incident as a security vulnerability, but Altman later changed the characterization, calling it a more fundamental "alignment failure" —

The model should never have cheated, it is OpenAI that failed to train it properly.

From now on, any alignment failure should be treated as a major event. We will take as much time as we need to figure it out. Making AI safe is more important than the development momentum of any company.

This statement will directly affect Astra. It will not be shelved, but new safety tests must be fully implemented before its release; as for when that will be, OpenAI executives have not given a timetable.

In an industry where the leading advantage is calculated by days, suspending the training of cutting-edge models may incur huge commercial costs.

Is this a real "reboot" or just commercial rhetoric?

*Time* calls OpenAI's recent round of changes a "reboot".

The product line is rebooted. OpenAI admits that it missed the programming and enterprise markets, decides to re-concentrate resources on Codex, use Agents to transform ChatGPT, and let engineer-turned Brockman take over the process of "turning research into revenue".

Its identity has also been rebooted. This company, which has lost members of its safety team repeatedly in the past few years and has been frequently criticized for putting commercialization above safety, is now trying to prove that in the face of risks, they are willing to slow down voluntarily for the benefit of society.

It seems like a story of rising up to reform. But the *Time* reporter also reminds us that this turn to safety may not be just a simple moral choice —