HomeArticle

550 Billion Zhipu AI, Refuses to Become a "Wage Earner" for Big Tech Giants

中国企业家杂志2026-09-02 08:28
Rejecting the mindless stacking of parameters, the large model competition has switched to new evaluation metrics.

Stop mindlessly stacking parameters, the large model competition has adopted a new evaluation metric.

On August 31, Zhipu AI released its 2026 semi-annual report. The financial report shows that from January to June, Zhipu AI's revenue reached 9539 million yuan, a year-on-year surge of about 400%; gross profit was 252 million yuan, up 163.7% year on year, with a loss of 2.072 billion yuan during the period, compared with a loss of 2.358 billion yuan in the same period of last year; R&D expenditure was 2.13 billion yuan, up 33% year on year.

Liu Debing, Chairman of Zhipu AI, said: If the key word for Zhipu AI last year was "upper limit of intelligence", the key word for the first half of this year is "capability delivery".

In terms of revenue structure, Zhipu AI's commercialization has entered a large-scale phase of rapid growth. In the first half of the year, revenue from the open platform and API business was about 825 million yuan, up 2735.7% year on year, and its proportion in total revenue rose from 15.2% in the same period of last year to 86.5%. This shows that Zhipu AI's revenue structure has undergone fundamental changes.

Although Zhipu AI's revenue growth is already very fast, it still falls short of the 1.3 billion yuan revenue expectation given by external institutions. The management also released another guidance signal: by the end of August 2026, the monthly measured ARR (Annual Recurring Revenue) has reached 1.6 billion US dollars. In August alone, the revenue hit 133 million US dollars, which is basically equivalent to the revenue scale of the first half of the year.

"For overseas cutting-edge model companies and another company in the industry, we multiply the data of the latest week by 52. Considering the volume release of GLM-5.3, if calculated according to this caliber, the latest ARR has exceeded 2 billion US dollars." said Xiao Lei, Secretary of the Board of Zhipu AI.

While revenue is skyrocketing, Zhipu AI's computing power expenditure remains heavy. The company's gross profit margin dropped from 50% in the same period of last year to 26.4%. At the same time, the price war of large models is still ongoing. GLM-5.3 was launched on August 19, and the price of the GLM-5.3-Flash version is even lower than that of "the King of Involution" DeepSeek-V4-Flash.

According to the AI industry research released by Barclays on August 28: For every 100 US dollars of revenue earned by model companies, about 35 to 40 US dollars will flow to the three major cloud giants in the form of reasoning computing power fees: Amazon AWS, Microsoft Azure and Google Cloud Platform.

In the financial report and the conference call, Zhipu AI's management emphasized the "North Star" indicator for the company to compete in the large model competition: leading in short-term list "benchmark scores" is no longer the top priority. The key to winning in the industry lies in who can continuously launch more powerful models to the market at a faster pace and lower cost.

To this end, Zhipu AI drew a diagram, in addition to its own models, it also selected manufacturers such as Kimi, DeepSeek, OpenAI and Anthropic for comparison. The horizontal axis represents paradigm progress, the vertical axis represents intelligence index, and the depth dimension represents cost per task.

Source: Zhipu AI 2026 Semi-Annual Report

Price remains a sensitive metric to measure the cost of large models. The diagram shows that before mid-August, DeepSeek ranked first in the intelligence-cost chart, but after DeepSeek-V4-Pro announced a price increase in mid-August, GLM-5.3 and GLM-5.3-Flash versions replaced DeepSeek as the new leader.

What supports Zhipu AI's overtaking is its firm technical path. Zhipu AI's management believes that with the same architecture, total parameters and activation parameters, only by expanding the post-training scale, the end-to-end completion rate can be increased by more than 50%. The performance of GLM-5.3 also confirms that Scaling (model size scaling) does not only have the path of stacking parameters, and post-training is the link with the greatest potential for capability improvement at the current stage.

However, this also raises external questions: when DeepSeek and Kimi are both sprinting for larger training parameters, Zhipu AI's GLM-5.3 series still uses the GLM-5.0 as its training base. Can simply increasing post-training investment and optimizing infrastructure help the model reach a higher upper limit of intelligence?

After the Hong Kong stock market opened on September 1, Zhipu AI's share price rose by nearly 5% at one point, then fluctuated and fell back in the afternoon. As of the close, Zhipu AI's share price was 1179 Hong Kong dollars, a slight drop of 1.34%, with a total market value of 550 billion Hong Kong dollars.

Model Companies Are Becoming More Asset-heavy

Zhipu AI is getting more and more asset-heavy. According to the financial report, it has 981 employees, twice as many as MiniMax. In addition, MiniMax made no acquisition or merger moves in the first half of the year, while Zhipu AI has made frequent capital moves.

The financial report disclosed the progress of two acquisitions of Zhipu AI: First, in May this year, Zhipu AI acquired 100% equity of Beijing Hongzuan at a cost of no more than 361 million yuan, whose main asset is the Diamond Building at Dongbeiwang West Road, Haidian District, Beijing. The other is that in June this year, Zhipu AI acquired 60% equity of Zhongke Jiahe, an AI infrastructure R&D manufacturer, for 292 million yuan in cash, and the acquisition was completed in June.

In July this year, Zhipu AI completed the placement of new shares (additional issuance) of 31.41 billion Hong Kong dollars. According to the announcement, the raised funds are mainly used for basic model R&D, computing power infrastructure, mergers and acquisitions, etc.

Self-built computing centers have become the core sector of capital expenditure for model companies. Some media reported that Zhipu AI has launched a 1GW-level data center, all using domestic chips, aiming to meet the needs of business expansion while seeking further reduction of computing power costs.

Source: AI Generated

Zhipu AI disclosed that the reasoning of GLM‑5.3‑Flash is provided by a domestic chip cluster. On the training and reasoning side, after realizing large-scale low-cost reasoning with 100,000-level domestic chips, the unit Token reasoning cost has decreased by 80% compared with the beginning of the year.

Computing power is becoming a bottleneck. Relying solely on leasing can no longer meet the requirements of model training and the explosive demand for reasoning. Models must be tightly coupled with infrastructure to achieve optimal efficiency and maximum capability.

According to Zhipu AI's financial report, from the cost side, the company's cost of sales increased from 95.45 million yuan to 702 million yuan, up 635.4% year on year, which is mainly affected by the increase in computing service fees.

The expenditure of model manufacturers is flowing to cloud vendors. In the second quarter of 2026, Alibaba Cloud's quarterly revenue was 48.437 billion yuan, corresponding to (Alibaba's first fiscal quarter of fiscal year 2027) a year-on-year increase of 45%, hitting the highest growth rate in 22 quarters.

In order to solve the challenges of memory capacity and bandwidth of domestic chips, Zhipu AI optimized the infrastructure layer to increase the density of intelligence per unit — "computing power multiplier" (the revenue from open platform and API corresponding to every 1 yuan of computing power invested by large model companies). Zhipu AI's computing power multiplier this year has increased by 14 times compared with the first half of last year, which also indicates that the company's computing power supply efficiency has been significantly improved.

How to Win Through Post-Training?

Zhipu AI's summary of its technical route is to use the right data to train at the right scale at the right time, so as to get the right model, not the largest one. The company pursues to obtain a model with sufficient intelligence benefits after reaching sufficient training depth under the current constraints of data and computing power.

This idea is also significantly different from that of Kimi and DeepSeek.

In mid-July, Kimi K3 was released, with a total parameter count of 2.8 trillion and activation parameters of about 104B, making it the open-source model with the largest parameter scale in the world at present. DeepSeek-V4-Pro released in April this year has a total parameter count of 1.6 trillion, a pre-training data volume of 33T, and activation parameters of about 49B.

However, GLM-5.0 launched by Zhipu AI in February this year has a parameter scale of only 744 billion, 40B activation, and a pre-training data volume of 28.5T, which lags behind Moonshot AI and DeepSeek in both parameter scale and pre-training.

In response to external questions, Tang Jie, founder of Zhipu AI, said that model size is a comprehensive choice that requires simultaneous consideration of resources, data and user availability. The domestic training data is usually in the order of 30T to 50T Tokens. Under this data scale, if the model is greatly expanded according to the Scaling Law, the marginal benefit may not be high enough; on the premise of limited computing power, it is more necessary to judge what scale is optimal.

"On the other hand, the potential of mid-training and post-training has not been fully released. At the current stage, simply expanding parameters is not the most critical. The company has not stopped scaling the base model, but is advancing in stages according to the established route." Tang Jie said.

The established route mentioned by Tang Jie refers to: In January 2026, the company trained GLM-5.3 using a base of about 744 billion parameters, and Zhipu AI still plans to use this base for more than half a year; the mid-training, post-training, task environment and SFT data arrangements for GLM-5.1, 5.2 and 5.3 were all planned at the beginning of the year.

"The pre-research of the reasoning system for the next-generation model has already started, with the goal of serving users at a sufficient scale on the first day of the model's release, rather than making a model with very large parameters that cannot be deployed economically." Tang Jie said.

Take GLM-5.3-Flash as an example. It adopts a brand-new architecture designed specifically to reduce costs. Its benchmark test and practical application performance are better than GLM-5.2, but its price is only one-tenth of that of GLM-5.2.

In addition, Tang Jie also mentioned the RSI (Recursive Self-Improvement) of the model, which he calls Fully Self-Training, allowing the model to participate in pre-training, mid-training and post-training, and gradually form self-evolution capabilities.

He said frankly: The biggest difficulty is not to continue expanding the model, but to enable the model to reliably judge when to stop, when to correct errors, and how to evaluate its own training results.

This article is from the WeChat Official Account "China Entrepreneur Magazine" (ID: iceo-com-cn), written by Li Yuan, edited by He Yifan, authorized for release by 36Kr.