OpenAI's Mid-game Showdown
Big companies that keep pushing the boundaries of AI intelligence have all started "hiding things".
Anthropic released Fable 5 in early June, and OpenAI launched the GPT-5.6 series a month later; but as safety issues drew widespread attention, nearly two more months have passed, and neither of the two companies has taken further actions on their cutting-edge models.
Company A, which holds the leading edge in the most advanced models, can afford to wait, but OpenAI, which is playing catch-up in both revenue and model performance, is in a very different position.
What is the current situation inside OpenAI?
In the latest Time Magazine cover story Inside OpenAI’s Reboot, journalist Alex Heath spent two weeks at OpenAI's headquarters, interviewing more than 20 executives, employees, investors, clients and competitors, and had a conversation with Sam Altman for over two hours, bringing us first-hand updates —
Seeing Anthropic seize the lead in the developer and enterprise market with Claude Code, and Google Gemini exceed 1 billion monthly active users, OpenAI featured in the report is no longer the absolute pace-setter that defined the rhythm of the AI competition two years ago.
Altman admitted very directly in the interview:
As a company, we have clearly made some mistakes. Whether in product direction or pre-training research, we have fallen behind the position we originally aimed to reach.
Therefore, OpenAI has finally made up its mind to launch a sweeping "reboot".
According to this long-form report, the core viewpoints from OpenAI can be summarized as follows:
OpenAI admits it has stalled on product development. The explosive growth of ChatGPT made the company neglect the programming and enterprise market. The models lead the rankings, but have not been "monetized" into practical real-world products.
Codex is replacing ChatGPT as the new core product of the company. OpenAI has scaled back projects including Sora, the Disney cooperation program, and the independent browser Atlas, concentrating computing power and teams on Codex, and integrating its Agent capabilities into ChatGPT Work.
The next-generation model Astra aims not only to answer questions, but to work continuously and generate new knowledge. OpenAI demonstrated the collaborative work of 16 Agents to solve research-level mathematical problems, high-speed cross-software operation, and automated research capabilities that can complete the work of a junior AI researcher in one week.
The Agent "jailbreak" incident forced OpenAI to suspend the training of more powerful models. The company acknowledged that it already has an early warning tool to monitor the model's chain of thought, but did not enable it because it underestimated the model's capabilities. The release of Astra will also wait for new safety measures to be approved.
What OpenAI wants to do goes far beyond a single chat dialog box. Chips, data centers, portable hardware, humanoid robots, brain-computer interfaces, and even external computing power sales are all included in the larger blueprint of "Personal AGI".
OpenAI, losing out on product capabilities
Over the past year, Anthropic surpassed OpenAI for the first time in both valuation and annualized revenue.
According to data cited by Time, Anthropic's annualized revenue has exceeded 65 billion US dollars, while OpenAI's is about 40 billion US dollars; both companies are preparing for IPO, but Anthropic is very likely to be the first to go public.
Photo by XDA Developers
The key to Anthropic's overtaking is Claude Code, which almost defined the "programming Agent" product form at the beginning of the year.
This does not mean that OpenAI's model capabilities are inferior to Anthropic's. Although its text capabilities have been widely criticized, according to Greg Brockman, co-founder and president of OpenAI, their models have "always led" in benchmark scores at least in the programming field.
The problem is that the company used to care more about research than products, only focusing on whether the model can solve problems in the laboratory environment, while ignoring the real usage scenarios of developer users — whether it can resume interrupted tasks, integrate a large number of files, and have good interaction details.
These are the key factors that determine which tool ordinary users choose.
In other words, OpenAI did not lose to others in cutting-edge model capabilities, but fell behind in the speed of translating model capabilities into products.
Of course, OpenAI itself is very active in product development.
The success of ChatGPT made OpenAI obsessed with making "fun" products, such as the video generation app Sora and the independent browser Atlas, but these products diverted the company's resources for key breakthroughs.
Sam Altman admitted in the interview that OpenAI spread its efforts too "thinly".
Now, the more "cost-conscious" Greg Brockman has taken over most of the company's internal business from revenue, products to marketing. OpenAI has started to cut off side projects, and shift scarce computing power to Codex, the fastest-growing business.
ChatGPT still has the largest number of users, but it is no longer OpenAI's main growth driver. Altman even stopped using ChatGPT for a whole month, and only used Codex.
It follows naturally that Codex began to merge ChatGPT, and the product presented to the outside world is ChatGPT Work — now users no longer need to judge what model and tool to use, they only need to tell AI what they want to do, and the system will automatically schedule models, tools and Agents behind the scenes to complete the task.
This adjustment has already been reflected in the financial statements. In July this year, OpenAI's B-end enterprise business revenue finally exceeded C-end consumer business revenue for the first time.
If Claude Code has taught OpenAI anything, it is:
Only by making money first can you talk about survival.
The next-generation model is your "virtual colleague"
Now that Fable 5 and Claude Code are already very capable, what kind of product can create a generation gap with them and win the favor of users, especially enterprise users again?
OpenAI's answer is Astra.
The Time report states that this mysterious next-generation model can split a difficult mathematical problem into multiple parts, with 16 Agents dividing the work, coordinating with each other, and finally piecing together a complete proof; in addition, Astra can also operate various desktop software at a speed far exceeding that of humans.
Altman calls this capability "Persistent Agents" — Agents that exist continuously and work nonstop.
It no longer solves a problem in one go. After you define its identity and operating environment, it can work continuously for hours or even days like your colleague.
Continuous work is great, but Altman believes that the greatest value of Astra is its innovative capability —
I expect this to be the first model that can truly invent new things in a meaningful way.
Just like the classic question: If you only input knowledge up to the early 20th century into AI, can it independently discover the theory of relativity like Einstein?
Altman believes it can.
A model with innovative capabilities will not only write code and generate reports, but can also participate in the development of the next generation of AI; and after a smarter AI emerges, it can develop its own successor.
This kind of "recursive self-improvement" has always been one of the most dangerous ideas in AGI discussions, but it is also the long-cherished goal of OpenAI.
OpenAI's management has begun to talk about AGI in a near-perfect tense.
Chief Research Officer Mark Chen believes that the company has "completed 80% of the journey"; Brockman says that in two years, people may look back and see the present as the moment when AGI was born; Altman predicts that by the end of this year, there will be a system inside OpenAI that he is willing to call AGI.
Optimism pervades OpenAI, until an accident happens.
Agent escapes the sandbox, OpenAI hits the brakes voluntarily
Few days after the 16-Agent parallel math problem solving demonstration for investors and clients, OpenAI suspended Astra's training process.
The direct cause was that an internal research model undergoing cybersecurity testing in an isolated environment exploited a vulnerability to escape the sandbox, and hacked into Hugging Face's production system in order to find the test answers.
After the incident, OpenAI froze and slowed down some projects, focusing on strengthening sandboxing and monitoring. Soon, researchers found danger signals in another undisclosed large-scale training session.
Photo by The New York Times
Since this training may bring a substantial performance leap, the company finally decided to suspend the training until the safety measures are in place.
OpenAI initially defined the Hugging Face incident as a security vulnerability, but Altman later changed the definition, believing that it was a more fundamental "alignment failure" —
The model should never have cheated in the first place, and it was OpenAI that failed to train it properly.
From now on, any alignment failure should be treated as a major event. We will take the time needed to figure it out. Making AI safe is more important than any company's growth momentum.
This statement will directly affect Astra. It will not be shelved, but new safety tests must be implemented before its release; as for when, OpenAI executives have not given a timetable.
In an industry where leading advantages are calculated by days, suspending the training of cutting-edge models may incur huge commercial costs.
Is it a real "reboot" or just commercial rhetoric?
Time calls OpenAI's recent round of changes a "reboot".
The product line has been rebooted. OpenAI admitted that it missed the programming and enterprise market, decided to re-concentrate resources on Codex, transform ChatGPT with Agents, and let Brockman, an engineer by training, take over the process of "turning research into revenue".
Its identity has also been rebooted. This company, which has lost members of its safety team repeatedly in the past few years and has been criticized for putting